
Product
Introducing Socket Scanning for VS Code Marketplace Extensions
Socket now scans VS Code extensions, giving teams early detection of risky behaviors, hidden capabilities, and supply chain threats in developer tools.
semantic-code-mcp
Advanced tools
AI-powered semantic code search for coding agents. MCP server with multi-provider embeddings and hybrid search.
AI-powered semantic code search for coding agents. An MCP server that indexes your codebase with vector embeddings so AI assistants can find code by meaning, not just keywords.
Ask "where do we handle authentication?" and find code that uses
login,session,verifyCredentials— even when no file contains the word "authentication."
Traditional grep and keyword search break down when you don't know the exact terms used in the codebase. Semantic search bridges that gap:
"error handling" finds try/catch, onRejected, fallback patterns"embeding modle" still finds embedding model codeBased on Cursor's research showing semantic search improves AI agent performance by 12.5%.
npm install -g semantic-code-mcp
Add to your MCP config:
{
"mcpServers": {
"semantic-code-mcp": {
"command": "semantic-code-mcp",
"args": ["--workspace", "/path/to/your/project"]
}
}
}
That's it. Your AI assistant now has semantic code search.
| Provider | Model | Privacy | Speed |
|---|---|---|---|
| Local (default) | nomic-embed-text-v1.5 | 100% local | ~50ms/chunk |
| Gemini | gemini-embedding-001 | API call | Fast, batched |
| OpenAI | text-embedding-3-small | API call | Fast |
| OpenAI-compatible | Any compatible endpoint | Varies | Varies |
| Vertex AI | Google Cloud models | GCP | Fast |
.smart-coding-cache/embeddings.dbThree modes to match your codebase:
smart (default) — regex-based, language-aware splittingast — Tree-sitter parsing for precise function/class boundariesline — simple fixed-size line chunksCPU capped at 50% during indexing. Your machine stays responsive.
| Tool | Description |
|---|---|
a_semantic_search | Find code by meaning. Hybrid semantic + exact match scoring. |
b_index_codebase | Trigger manual reindex (normally automatic & incremental). |
c_clear_cache | Reset embeddings cache entirely. |
d_check_last_version | Look up latest package version from 20+ registries. |
e_set_workspace | Switch project at runtime without restart. |
f_get_status | Server health: version, index progress, config. |
| IDE / App | Guide | ${workspaceFolder} |
|---|---|---|
| VS Code | Setup | ✅ |
| Cursor | Setup | ✅ |
| Windsurf | Setup | ❌ |
| Claude Desktop | Setup | ❌ |
| OpenCode | Setup | ❌ |
| Raycast | Setup | ❌ |
| Antigravity | Setup | ❌ |
{
"mcpServers": {
"code-frontend": {
"command": "semantic-code-mcp",
"args": ["--workspace", "/path/to/frontend"]
},
"code-backend": {
"command": "semantic-code-mcp",
"args": ["--workspace", "/path/to/backend"]
}
}
}
All settings via environment variables. Prefix: SMART_CODING_.
| Variable | Default | Description |
|---|---|---|
SMART_CODING_VERBOSE | false | Detailed logging |
SMART_CODING_MAX_RESULTS | 5 | Search results returned |
SMART_CODING_BATCH_SIZE | 100 | Files per parallel batch |
SMART_CODING_MAX_FILE_SIZE | 1048576 | Max file size (1MB) |
SMART_CODING_CHUNK_SIZE | 25 | Lines per chunk |
SMART_CODING_CHUNKING_MODE | smart | smart / ast / line |
SMART_CODING_WATCH_FILES | false | Auto-reindex on changes |
SMART_CODING_AUTO_INDEX_DELAY | 5000 | Background index delay (ms) |
SMART_CODING_MAX_CPU_PERCENT | 50 | CPU cap during indexing |
| Variable | Default | Description |
|---|---|---|
SMART_CODING_EMBEDDING_PROVIDER | local | local / gemini / openai / openai-compatible / vertex |
SMART_CODING_EMBEDDING_MODEL | nomic-ai/nomic-embed-text-v1.5 | Model name |
SMART_CODING_EMBEDDING_DIMENSION | 128 | MRL dimension (64–768) |
SMART_CODING_DEVICE | auto | cpu / webgpu / auto |
| Variable | Default | Description |
|---|---|---|
SMART_CODING_GEMINI_API_KEY | — | API key |
SMART_CODING_GEMINI_MODEL | gemini-embedding-001 | Model |
SMART_CODING_GEMINI_DIMENSIONS | 768 | Output dimensions |
SMART_CODING_GEMINI_BATCH_SIZE | 24 | Micro-batch size |
SMART_CODING_GEMINI_MAX_RETRIES | 3 | Retry count |
| Variable | Default | Description |
|---|---|---|
SMART_CODING_EMBEDDING_API_KEY | — | API key |
SMART_CODING_EMBEDDING_BASE_URL | — | Base URL (compatible only) |
| Variable | Default | Description |
|---|---|---|
SMART_CODING_VERTEX_PROJECT | — | GCP project ID |
SMART_CODING_VERTEX_LOCATION | us-central1 | Region |
| Variable | Default | Description |
|---|---|---|
SMART_CODING_VECTOR_STORE_PROVIDER | sqlite | sqlite / milvus |
SMART_CODING_MILVUS_ADDRESS | — | Milvus endpoint |
SMART_CODING_MILVUS_TOKEN | — | Auth token |
SMART_CODING_MILVUS_DATABASE | default | Database name |
SMART_CODING_MILVUS_COLLECTION | smart_coding_embeddings | Collection |
| Variable | Default | Description |
|---|---|---|
SMART_CODING_SEMANTIC_WEIGHT | 0.7 | Semantic vs exact weight |
SMART_CODING_EXACT_MATCH_BOOST | 1.5 | Exact match multiplier |
{
"mcpServers": {
"semantic-code-mcp": {
"command": "semantic-code-mcp",
"args": ["--workspace", "/path/to/project"],
"env": {
"SMART_CODING_EMBEDDING_PROVIDER": "gemini",
"SMART_CODING_GEMINI_API_KEY": "YOUR_KEY",
"SMART_CODING_VECTOR_STORE_PROVIDER": "milvus",
"SMART_CODING_MILVUS_ADDRESS": "http://localhost:19530"
}
}
}
}
semantic-code-mcp/
├── index.js # MCP server entry point
├── lib/
│ ├── config.js # Configuration loader
│ ├── cache-factory.js # SQLite / Milvus provider selection
│ ├── cache.js # SQLite vector store
│ ├── milvus-cache.js # Milvus vector store
│ ├── mrl-embedder.js # Local MRL embedder
│ ├── gemini-embedder.js# Gemini API embedder
│ ├── ast-chunker.js # Tree-sitter AST chunking
│ ├── tokenizer.js # Token counting
│ └── utils.js # Cosine similarity, hashing, smart chunking
├── features/
│ ├── hybrid-search.js # Semantic + exact match search
│ ├── index-codebase.js # File discovery & incremental indexing
│ ├── clear-cache.js # Cache reset
│ ├── check-last-version.js # Package version lookup
│ ├── set-workspace.js # Runtime workspace switching
│ └── get-status.js # Server status
└── test/ # Vitest test suite
Your code files
↓ glob + .gitignore-aware discovery
Smart/AST chunking
↓ language-aware splitting
AI embedding (local or API)
↓ vector generation
SQLite or Milvus storage
↓ incremental, hash-based updates
Search query
↓ embed query → cosine similarity → exact match boost
Top N results with relevance scores
Progressive indexing — search works immediately while indexing continues in the background. Only changed files are re-indexed on subsequent runs.
MIT License
Copyright (c) 2025 Omar Haris (original), bitkyc08 (modifications, 2026)
See LICENSE for full text.
Built on smart-coding-mcp by Omar Haris. Extended with multi-provider embeddings, Milvus ANN search, AST chunking, resource throttling, and comprehensive test suite.
FAQs
AI-powered semantic code search for coding agents. MCP server with multi-provider embeddings and hybrid search.
The npm package semantic-code-mcp receives a total of 14 weekly downloads. As such, semantic-code-mcp popularity was classified as not popular.
We found that semantic-code-mcp demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Product
Socket now scans VS Code extensions, giving teams early detection of risky behaviors, hidden capabilities, and supply chain threats in developer tools.

Research
/Security News
Socket uncovered two malicious VS Code themes in a GlassWorm-linked cluster with thousands of installs across VS Code Marketplace and Open VSX.

Security News
/Company News
Capital One is partnering with Socket to proactively secure its open source supply chain.