
Security News
GPT-6 Astra Attempts Supply Chain Attacks Against Open Source Maintainers in Testing
GPT-6 Astra hits 100% on ExploitBench and finds zero-days autonomously, while independent tests reveal scope violations and monitoring gaps.
opencode-codebase-index
Advanced tools
Semantic codebase indexing and search for OpenCode - find code by meaning, not just keywords
Stop grepping for concepts. Start searching for meaning.
opencode-codebase-index brings semantic understanding to your OpenCode, Pi, Codex, and MCP-compatible workflows. Instead of guessing function names or grepping for keywords, ask your codebase questions in plain English.
check_creds.tree-sitter and usearch. Incremental updates take milliseconds.Install the plugin
npm install opencode-codebase-index
Add to opencode.json
{
"plugin": ["opencode-codebase-index"]
}
Index your codebase
Run /index or ask the agent to index your codebase. This only needs to be done once — subsequent updates are incremental.
Recommended check: run /status after the first index so you can confirm the detected provider/model before you start searching.
Start Searching Ask:
"Find the function that handles credit card validation errors"
Install as a Pi package to get first-class codebase_search, index_codebase, call graph, PR impact, and knowledge-base tools plus the codebase-search skill.
pi install npm:opencode-codebase-index
# or, for local development:
pi install ./path/to/opencode-codebase-index
Pi uses the neutral .codebase-index/ project storage and falls back to existing OpenCode state when present.
Install once for Codex threads and get skill guidance plus MCP tools in one manifest.
codex plugin marketplace add Helweg/opencode-codebase-index
codex plugin add codebase-index@helweg-plugins
index_codebase, index_status, codebase_search, etc.) and the codebase-search skill guidance.The plugin includes:
skills/ guidance for local workflowshooks/hooks.json lightweight session-start guidance.mcp.json running the published opencode-codebase-index CLI via npx … --host codex, so a git marketplace install works without a local build.agents/plugins/marketplace.json so this repo can act as a Codex marketplace sourceInstall once for Claude Code sessions and get skill guidance plus MCP tools in one manifest.
/plugin marketplace add Helweg/opencode-codebase-index
/plugin install codebase-index@helweg-plugins
index_codebase, index_status, codebase_search, etc.) and the codebase-search skill guidance.The plugin includes:
skills/ guidance for local workflowshooks/hooks.json lightweight session-start guidancemcpServers (in .claude-plugin/plugin.json) running the published opencode-codebase-index CLI via npx … --host claude, so a git marketplace install works without a local build.claude-plugin/marketplace.json so this repo can act as a Claude Code marketplace sourceDefault auto-detect order: Ollama → GitHub Copilot → OpenAI → Google
Ollama is the preferred zero-cost local option and works especially well for large repos:
ollama pull nomic-embed-text
{
"embeddingProvider": "ollama"
}
GitHub Copilot is a good default if OpenCode already has Copilot auth and you prefer hosted embeddings.
OpenAI is a good hosted option when you want predictable API behavior and standard cloud setup.
Google is available if you prefer Gemini-hosted embeddings.
If /status reports provider or compatibility problems, follow that guidance before using /index force.
Use the same semantic search from any MCP-compatible client. Index once, search from anywhere.
Install dependencies
npm install opencode-codebase-index @modelcontextprotocol/sdk zod
Configure your MCP client
Cursor (.cursor/mcp.json):
{
"mcpServers": {
"codebase-index": {
"command": "npx",
"args": ["opencode-codebase-index-mcp", "--project", "/path/to/your/project"]
}
}
}
Claude Code (claude_desktop_config.json):
{
"mcpServers": {
"codebase-index": {
"command": "npx",
"args": ["opencode-codebase-index-mcp", "--project", "/path/to/your/project"]
}
}
}
CLI options
npx opencode-codebase-index-mcp --project /path/to/repo # specify project root
npx opencode-codebase-index-mcp --config /path/to/config # custom config file
npx opencode-codebase-index-mcp # uses current directory
The MCP server exposes all 12 tools (codebase_search, codebase_peek, find_similar, implementation_lookup, call_graph, call_graph_path, pr_impact, index_codebase, index_status, index_health_check, index_metrics, index_logs) and 5 prompts (search, find, definition, index, status).
The MCP dependencies (@modelcontextprotocol/sdk, zod) are optional peer dependencies — they're only needed if you use the MCP server.
Scenario: You're new to a codebase and need to fix a bug in the payment flow.
Without Plugin (grep):
grep "payment" . → 500 results (too many)grep "card" . → 200 results (mostly UI)grep "stripe" . → 50 results (maybe?)With opencode-codebase-index:
You ask: "Where is the payment validation logic?"
Plugin returns:
src/services/billing.ts:45 (Class PaymentValidator)
src/utils/stripe.ts:12 (Function validateCardToken)
src/api/checkout.ts:89 (Route handler for /pay)
| Scenario | Tool | Why |
|---|---|---|
| Don't know the function name | codebase_search | Semantic search finds by meaning |
| Exploring unfamiliar codebase | codebase_search | Discovers related code across files |
| Just need to find locations | codebase_peek | Returns metadata only, saves ~90% tokens |
| Need the authoritative definition site | implementation_lookup | Prioritizes real implementation definitions over docs/tests |
| Understand code flow | call_graph | Find callers/callees of any function |
| Trace dependency paths | call_graph_path | Find the shortest known call path between two symbols |
| Know exact identifier | grep | Faster, finds all occurrences |
| Need ALL matches | grep | Semantic returns top N only |
| Mixed discovery + precision | /find (hybrid) | Best of both worlds |
Rule of thumb: codebase_peek to find locations → Read to examine → grep for precision. For symbol-definition questions, use implementation_lookup first.
Recent OMO releases include a built-in CodeGraph MCP and make it part of the default agent workflow. This does not replace opencode-codebase-index; the two tools answer different first questions.
| Need | Prefer | Why |
|---|---|---|
| Find code by intent, behavior, or natural language | codebase_peek / codebase_search | Semantic + hybrid retrieval works when you do not know exact names |
| Jump to the likely implementation site | implementation_lookup | Definition-oriented ranking prefers source over tests/docs |
| Find similar implementations or duplicate patterns | find_similar | Embedding similarity compares code shape and meaning |
| Follow callers, callees, imports, inheritance, or implementations | OMO CodeGraph or call_graph | Structural graph tools are best for dependency topology |
| Find a shortest known relationship chain | call_graph_path | Uses this plugin's indexed call edges to connect two symbols |
| Include external docs, examples, or API references in discovery | add_knowledge_base + codebase_search | Knowledge bases are indexed into the same retrieval store |
Recommended OMO workflow:
codebase_peek when the prompt is conceptual, such as "where is auth enforced?" or "payment validation flow".implementation_lookup once you have a symbol or concept that should resolve to a definition.call_graph, or call_graph_path after locating the relevant symbol to check blast radius and dependency flow.grep for exact identifiers and exhaustive text matches.If OMO reports an uninitialized CodeGraph workspace, follow its codegraph init guidance. That setup is independent from this plugin's index under .opencode/index/, so /index and codegraph init may both be useful in the same repository.
In our testing across open-source codebases (axios, express), we observed up to 90% reduction in token usage for conceptual queries like "find the error handling middleware".
graph TD
subgraph Indexing
A[Source Code] -->|Tree-sitter| B[Semantic Chunks]
B -->|Embedding Model| C[Vectors]
C -->|uSearch| D[(Vector Store)]
C -->|SQLite| G[(Embeddings DB)]
B -->|BM25| E[(Inverted Index)]
B -->|Branch Catalog| G
end
subgraph Searching
Q[User Query] -->|Embedding Model| V[Query Vector]
V -->|Cosine Similarity| D
Q -->|BM25| E
D --> F[Hybrid Fusion RRF/Weighted]
E --> F
F --> X[Deterministic Rerank]
G -->|Branch + Metadata Filters| X
X --> R[Ranked Results]
end
tree-sitter to intelligently parse your code into meaningful blocks (functions, classes, interfaces). JSDoc comments and docstrings are automatically included with their associated code.Supported Languages (Tree-sitter semantic parsing): TypeScript, JavaScript, Python, Rust, Go, Java, C#, Ruby, PHP, Apex, Bash, C, C++, JSON, TOML, YAML, Zig, GDScript, MATLAB†
† MATLAB (.m) is opt-in — see below.
Additional Supported Formats (line-based chunking): TXT, HTML, HTM, Markdown, Shell scripts
Default File Patterns:
**/*.{ts,tsx,js,jsx,mjs,cjs} **/*.{py,pyi}
**/*.{go,rs,java,kt,scala} **/*.{c,cpp,cc,h,hpp}
**/*.{rb,php,inc,swift} **/*.{vue,svelte,astro}
**/*.{sql,graphql,proto} **/*.{yaml,yml,toml}
**/*.{md,mdx} **/*.{sh,bash,zsh}
**/*.{txt,html,htm} **/*.{cls,trigger}
**/*.zig **/*.gd
Use include to replace defaults, or additionalInclude to extend (e.g. "**/*.pdf", "**/*.csv").
†MATLAB opt-in: .m is excluded from defaults because it conflicts with the Objective-C extension used on Apple codebases. To enable MATLAB discovery, add to your global config (~/.config/opencode/codebase-index.json):
{ "additionalInclude": ["**/*.m"] }
Max File Size: Default 1MB (1048576 bytes). Configure via indexing.maxFileSize (bytes).
2. Chunking: Large blocks are split with overlapping windows to preserve context across chunk boundaries.
3. Embedding: These blocks are converted into vector representations using your configured AI provider.
4. Storage: Embeddings are stored in SQLite (deduplicated by content hash) and vectors in usearch with F16 quantization for 50% memory savings. A branch catalog tracks which chunks exist on each branch.
5. Hybrid Search: Combines semantic similarity (vectors) with BM25 keyword matching, fuses (rrf default, weighted fallback), applies deterministic rerank, then filters by current branch/metadata.
Performance characteristics:
The plugin automatically detects git branches and optimizes indexing across branch switches.
When you switch branches, code changes but embeddings for unchanged content remain the same. The plugin:
| Scenario | Without Branch Awareness | With Branch Awareness |
|---|---|---|
| Switch to feature branch | Re-index everything | Instant — reuse existing embeddings |
| Return to main | Re-index everything | Instant — catalog already exists |
| Search on branch | May return stale results | Only returns current branch's code |
.git/HEAD.opencode/index/
├── codebase.db # SQLite: embeddings, chunks, branch catalog, symbols, call edges
├── vectors.usearch # Vector index (uSearch)
├── inverted-index.json # BM25 keyword index
└── file-hashes.json # File change detection
The following files/folders are excluded from indexing by default:
. (e.g., .github, .vscode, .env)build, mingwBuildDebug, cmake-build-debug)node_modules, dist, vendor, __pycache__, target, coverage, etc.The plugin exposes these tools to the OpenCode agent:
codebase_searchThe primary tool. Searches code by describing behavior.
"find the middleware that sanitizes input"search.fusionStrategy) → deterministic rerank (search.rerankTopN) → filtersWriting good queries:
| ✅ Good queries (describe behavior) | ❌ Bad queries (too vague) |
|---|---|
| "function that validates email format" | "email" |
| "error handling for failed API calls" | "error" |
| "middleware that checks authentication" | "auth middleware" |
| "code that calculates shipping costs" | "shipping" |
| "where user permissions are checked" | "permissions" |
codebase_peekToken-efficient discovery. Returns only metadata (file, line, name, type) without code content.
codebase_search.codebase_search (metadata-only output)[1] function "validatePayment" at src/billing.ts:45-67 (score: 0.92)
[2] class "PaymentProcessor" at src/processor.ts:12-89 (score: 0.87)
Use Read tool to examine specific files.
codebase_peek → find locations → Read specific filesimplementation_lookupDefinition-first lookup. Jumps to the authoritative definition site for a symbol or natural-language definition query.
codebase_search for broader discovery.find_similarFind code similar to a provided snippet.
index_codebaseManually trigger indexing.
force (rebuild all), estimateOnly (check costs), verbose (show skipped files and parse failures).index_statusChecks if the index is ready and healthy.
/index to confirm the detected provider/model and whether the index is ready to search.index_health_checkMaintenance tool to remove stale entries from deleted files and orphaned embeddings/chunks from the database.
index_metricsReturns collected metrics about indexing and search performance. Requires debug.enabled and debug.metrics to be true.
index_logsReturns recent debug logs with optional filtering.
category (optional: search, embedding, cache, gc, branch), level (optional: error, warn, info, debug), limit (default: 50).call_graphQuery the call graph to find callers or callees of a function/method. Automatically built during indexing for TypeScript, JavaScript, Python, Go, Rust, PHP, Apex, Zig, GDScript, and MATLAB.
name (function name), direction (callers or callees), symbolId (required for callees, returned by previous queries), relationshipType (optional: Call, MethodCall, Constructor, Import, Inherits, Implements).validateToken → call_graph(name="validateToken", direction="callers")call_graph_pathFind the shortest known call-graph path between two symbols. Use it after codebase_peek, implementation_lookup, or call_graph identifies the important source and target names.
from (source symbol name), to (target symbol name), maxDepth (optional, default 10).createOrder reaches chargeCard → call_graph_path(from="createOrder", to="chargeCard")pr_impactAnalyzes a PR's changed files to determine impact scope within the codebase.
checkConflicts (optional, default false) — when true, detects overlapping concurrent PRs sharing affected symbols and returns conflictingPRs.index_visualizeGenerate a self-contained temporal call graph view for browser-based exploration.
src/tools or native; symbols are indexed functions, classes, methods, or similar named code units; edges are caller/callee relationships.directory (optional folder filter), maxNodes (default 5000), includeOrphans (include disconnected symbols).index_visualize(directory="src/tools", maxNodes=1500)CLI shortcut after building locally:
npm run build
npm run visualize
npm run visualize -- native
npm run visualize -- src/tools max=1000
npm run visualize -- src/indexer orphans
add_knowledge_baseAdd a folder as a knowledge base to be indexed alongside project code.
path (folder path, absolute or relative), reindex (optional, default true)./etc, /proc, /sys, /dev) and sensitive home directories (.ssh, .gnupg, .aws, .docker, .kube) are blocked. Symlinks are resolved before validation.add_knowledge_base(path="/path/to/docs")list_knowledge_basesList all configured knowledge base folders and their status.
remove_knowledge_baseRemove a knowledge base folder from the index.
path (folder path to remove), reindex (optional, default false).remove_knowledge_base(path="/path/to/docs")The plugin automatically registers these slash commands:
| Command | Description |
|---|---|
/definition <query> | Definition Lookup. Finds the authoritative implementation site for a symbol or concept. |
/peek <query> | Quick Semantic Lookup. Returns likely locations only, without full code content. |
/reindex | Full Rebuild. Rebuilds the codebase index from scratch. |
/search <query> | Pure Semantic Search. Best for "How does X work?" |
/find <query> | Hybrid Search. Combines semantic search + grep. Best for "Find usage of X". |
/call-graph <query> | Call Graph Trace. Find callers/callees to understand execution flow. |
/pr-impact <PR number or branch> | PR Impact Analysis. Analyze changed files, affected symbols, communities, hub nodes, and risk. |
| `/visualize [directory | max=N |
/index | Update Index. Runs incremental indexing by default; use /index force for a full rebuild. |
/status | Check Status. Shows if indexed, chunk count, and provider info. |
The plugin can index external documentation alongside your project code. The indexed codebase includes:
Use the built-in tools to add documentation folders:
add_knowledge_base(path="/path/to/api-docs")
add_knowledge_base(path="/path/to/examples")
The folder will be indexed into the same database as your project code. All searches automatically include both sources.
list_knowledge_bases # Show configured knowledge bases
remove_knowledge_base(path="/path/to/api-docs") # Remove a knowledge base
Project-level config (.opencode/codebase-index.json):
{
"knowledgeBases": [
"/home/user/docs/esp-idf",
"/home/user/docs/arduino"
]
}
Global-level config (~/.config/opencode/codebase-index.json):
{
"embeddingProvider": "custom",
"customProvider": {
"baseUrl": "{env:EMBED_BASE_URL}",
"model": "BAAI/bge-m3",
"dimensions": 1024,
"apiKey": "{env:EMBED_API_KEY}"
}
}
Config merging: Global config is the base, project config overrides. Knowledge bases from both levels are merged.
/index force after changesThe plugin supports API-based reranking for improved search result quality. Reranking uses a cross-encoder model to rescore the top search results.
Add to your config (.opencode/codebase-index.json or global config):
{
"reranker": {
"enabled": true,
"baseUrl": "https://api.cohere.ai/v1",
"model": "rerank-v3.5",
"apiKey": "{env:RERANK_API_KEY}",
"topN": 20
}
}
| Option | Default | Description |
|---|---|---|
enabled | false | Enable reranking |
baseUrl | - | Rerank API endpoint |
model | - | Reranking model name |
apiKey | - | API key (use {env:VAR} for security) |
topN | 20 | Number of top results to rerank |
timeoutMs | 30000 | Request timeout |
Query → Embedding Search → BM25 Search → Fusion → Reranking → Results
Any OpenAI-compatible reranking endpoint. Examples:
BAAI/bge-reranker-v2-m3rerank-english-v3.0/v1/rerank formatOpenCode default (existing behavior):
.opencode/codebase-index.json.opencode/index~/.config/opencode/codebase-index.json~/.opencode/global-indexCodex/Pi host mode (neutral default):
.codebase-index/config.json.codebase-index/index~/.config/codebase-index/config.json~/.codebase-index/global-indexClaude Code host mode (--host claude):
.claude/codebase-index.json.claude/index~/.claude/codebase-index.json~/.claude/global-indexCodex, Claude Code, and Pi read legacy OpenCode paths when host-native paths are absent, so existing state continues to work.
Zero-config by default (uses auto mode). Customize in .opencode/codebase-index.json:
{
// === Embedding Provider ===
"embeddingProvider": "custom", // auto | github-copilot | openai | google | ollama | custom
"scope": "project", // project (per-repo) | global (shared)
// === Custom Embedding API (when embeddingProvider is "custom") ===
"customProvider": {
"baseUrl": "{env:EMBED_BASE_URL}",
"model": "BAAI/bge-m3",
"dimensions": 1024,
"apiKey": "{env:EMBED_API_KEY}",
"maxTokens": 8192, // Max tokens per input text
"timeoutMs": 30000, // Request timeout (ms)
"concurrency": 3, // Max concurrent requests
"requestIntervalMs": 1000, // Min delay between requests (ms)
"maxBatchSize": 64 // Max inputs per /embeddings request
},
// === File Patterns ===
"include": [ // Override default include patterns
"**/*.{ts,js,py,go,rs}"
],
"exclude": [ // Override default exclude patterns
"**/node_modules/**"
],
"additionalInclude": [ // Extend defaults (not replace)
"**/*.{txt,html,htm}",
"**/*.pdf"
],
// === Knowledge Bases ===
"knowledgeBases": [ // External docs to index alongside code
"/home/user/docs/esp-idf",
"/home/user/docs/arduino"
],
// === Indexing ===
"indexing": {
"autoIndex": false, // Auto-index on plugin load
"watchFiles": true, // Re-index on file changes
"maxFileSize": 1048576, // Max file size in bytes (default: 1MB)
"maxChunksPerFile": 100, // Max chunks per file
"semanticOnly": false, // Only index functions/classes (skip blocks)
"retries": 3, // Embedding API retry attempts
"retryDelayMs": 1000, // Delay between retries (ms)
"autoGc": true, // Auto garbage collection
"gcIntervalDays": 7, // GC interval (days)
"gcOrphanThreshold": 100, // GC trigger threshold
"requireProjectMarker": true, // Require .git/package.json to index
"maxDepth": 5, // Max directory depth (-1=unlimited, 0=root only)
"maxFilesPerDirectory": 100, // Max files per directory (smallest first)
"fallbackToTextOnMaxChunks": true // Fallback to text chunking on maxChunksPerFile
},
// === Search ===
"search": {
"maxResults": 20, // Max results to return
"minScore": 0.1, // Min similarity score (0-1)
"hybridWeight": 0.5, // Keyword (1.0) vs semantic (0.0)
"fusionStrategy": "rrf", // rrf | weighted
"rrfK": 60, // RRF smoothing constant
"rerankTopN": 20, // Deterministic rerank depth
"contextLines": 0, // Extra lines before/after match
"routingHints": true, // Runtime nudges for local discovery/definition queries
"routingGraphHandoffHints": false, // Add opt-in graph/OMO CodeGraph handoff wording
"routingHintRole": "system" // system | developer (message role used for hints)
},
"reranker": {
"enabled": false,
"provider": "cohere",
"model": "rerank-v3.5",
"apiKey": "{env:RERANK_API_KEY}",
"topN": 15,
"timeoutMs": 10000
},
"debug": {
"enabled": false, // Enable debug logging
"logLevel": "info", // error | warn | info | debug
"logSearch": true, // Log search operations
"logEmbedding": true, // Log embedding API calls
"logCache": true, // Log cache hits/misses
"logGc": true, // Log garbage collection
"logBranch": true, // Log branch detection
"metrics": false // Enable metrics collection
}
}
String values in codebase-index.json can reference environment variables with {env:VAR_NAME} when the placeholder is the entire string value. Variable names must match [A-Z_][A-Z0-9_]*. This is useful for secrets such as custom provider API keys so they do not need to be committed to the config file.
{
"embeddingProvider": "custom",
"customProvider": {
"baseUrl": "{env:EMBED_BASE_URL}",
"model": "nomic-embed-text",
"dimensions": 768,
"apiKey": "{env:EMBED_API_KEY}"
}
}
| Option | Default | Description |
|---|---|---|
embeddingProvider | "auto" | Which AI to use: auto, github-copilot, openai, google, ollama, custom |
scope | "project" | project = index per repo, global = shared index across repos |
include | (defaults) | Override the default include patterns (replaces defaults) |
exclude | (defaults) | Override the default exclude patterns (replaces defaults) |
additionalInclude | [] | Additional file patterns to include (extends defaults, e.g. "**/*.txt", "**/*.html") |
knowledgeBases | [] | External directories to index as knowledge bases (absolute or relative paths) |
| indexing | ||
autoIndex | false | Automatically index on plugin load |
watchFiles | true | Re-index when files change |
maxFileSize | 1048576 | Skip files larger than this (bytes). Default: 1MB |
maxChunksPerFile | 100 | Maximum chunks to index per file (controls token costs for large files) |
semanticOnly | false | When true, only index semantic nodes (functions, classes) and skip generic blocks |
retries | 3 | Number of retry attempts for failed embedding API calls |
retryDelayMs | 1000 | Delay between retries in milliseconds |
autoGc | true | Automatically run garbage collection to remove orphaned embeddings/chunks |
gcIntervalDays | 7 | Run GC on initialization if last GC was more than N days ago |
gcOrphanThreshold | 100 | Run GC after indexing if orphan count exceeds this threshold |
requireProjectMarker | true | Require a project marker (.git, package.json, etc.) to enable file watching and auto-indexing. Prevents accidentally indexing large directories like home. Set to false to index any directory. |
maxDepth | 5 | Max directory traversal depth. -1 = unlimited, 0 = only files in root dir, 1 = one level of subdirectories, etc. |
maxFilesPerDirectory | 100 | Max files to index per directory. Always picks the smallest files first. |
fallbackToTextOnMaxChunks | true | When a file exceeds maxChunksPerFile, fallback to text-based (line-by-line) chunking instead of skipping the rest of the file. |
| search | ||
maxResults | 20 | Maximum results to return |
minScore | 0.1 | Minimum similarity score (0-1). Lower = more results |
hybridWeight | 0.5 | Balance between keyword (1.0) and semantic (0.0) search |
fusionStrategy | "rrf" | Hybrid fusion mode: "rrf" (rank-based reciprocal rank fusion) or "weighted" (legacy score blending fallback) |
rrfK | 60 | RRF smoothing constant. Higher values flatten rank impact, lower values prioritize top-ranked candidates more strongly |
rerankTopN | 20 | Deterministic rerank depth cap. Applies lightweight name/path/chunk-type rerank to top-N only |
contextLines | 0 | Extra lines to include before/after each match |
routingHints | true | Inject lightweight runtime hints for local conceptual discovery and definition lookups. Set to false to disable plugin-side routing nudges. |
routingGraphHandoffHints | false | When true, conceptual discovery hints also say to use graph tools (including OMO CodeGraph) after semantic discovery identifies relevant symbols. |
routingHintRole | "system" | Message role used when injecting routing hints: "system" (default) or "developer". |
| reranker | Optional second-stage model reranker for the top candidate pool | |
enabled | false | Turn external reranking on/off |
provider | "custom" | Hosted shortcuts: cohere, jina, or custom |
model | — | Reranker model name required when enabled |
baseUrl | provider default | Override reranker endpoint base URL. cohere → https://api.cohere.ai/v1, jina → https://api.jina.ai/v1 |
apiKey | — | API key for hosted reranker providers |
topN | 15 | Number of top candidates to send to the external reranker |
timeoutMs | 10000 | Timeout for external rerank requests |
| debug | ||
enabled | false | Enable debug logging and metrics collection |
logLevel | "info" | Log level: error, warn, info, debug |
logSearch | true | Log search operations with timing breakdown |
logEmbedding | true | Log embedding API calls (success, error, rate-limit) |
logCache | true | Log cache hits and misses |
logGc | true | Log garbage collection operations |
logBranch | true | Log branch detection and switches |
metrics | false | Enable metrics collection (indexing stats, search timing, cache performance) |
When debug logging is enabled, the indexer now emits warn-level recovery messages if persisted cache state cannot be read safely.
file-hashes.json causes the in-memory file hash cache to be reset.failed-batches.json causes persisted retry batches to be skipped for that run.These warnings improve observability but do not change the recovery behavior: the indexer still falls back to a safe reset/skip path instead of crashing. If these warnings recur, remove the affected file under .opencode/index/ (or the global index directory) and rebuild with /index force.
codebase_search and codebase_peek use the hybrid path: semantic + keyword retrieval → fusion (fusionStrategy) → deterministic rerank (rerankTopN) → optional external reranker (reranker) → filtering.search.routingHints is enabled (default), the plugin adds tiny per-turn runtime hints for local conceptual discovery and definition queries. Conceptual discovery is nudged toward codebase_peek / codebase_search, while definition questions are nudged toward implementation_lookup. Exact identifier and unrelated operational tasks are left alone. Set search.routingGraphHandoffHints to true to add opt-in graph/OMO CodeGraph handoff wording, and set search.routingHintRole to "developer" if your client/runtime expects developer-role guidance instead of system-role guidance.find_similar stays semantic-only: semantic retrieval + deterministic rerank only (no keyword retrieval, no RRF).search.fusionStrategy to "weighted" to use the legacy weighted fusion path.benchmarks/baselines/retrieval-baseline.jsonbenchmark-results/retrieval-candidate.jsonThis repository includes a first-class eval system for retrieval quality with versioned golden sets, compare mode, parameter sweeps, CI budgets, and run artifacts.
npm run eval
npm run eval:ci
npm run eval:ci:ollama
npm run eval:compare -- --against benchmarks/baselines/eval-baseline-summary.json
CI usage split:
npm run eval:smoke: harness smoke check with local mock embeddings (used in main CI)npm run eval:ci: real quality gate against baseline/budget (for scheduled/manual quality workflow)For eval-quality.yml, the default CI path uses GitHub Models with the workflow GITHUB_TOKEN plus models: read, so you do not need a separate OpenAI API key just to run the scheduled gate.
That default GitHub Models path uses benchmarks/budgets/github-models.json, which applies stable absolute thresholds instead of the stricter baseline-regression budget used for explicit external providers.
Optional override secrets for another OpenAI-compatible endpoint:
EVAL_EMBED_BASE_URLEVAL_EMBED_API_KEYEVAL_EMBED_MODEL (optional, default text-embedding-3-small)EVAL_EMBED_DIMENSIONS (optional, default 1536)If you override the provider, set both EVAL_EMBED_BASE_URL and EVAL_EMBED_API_KEY. Otherwise the workflow falls back to GitHub Models automatically. Override providers continue to use the baseline-driven budget in benchmarks/budgets/default.json.
No OpenAI API access? Use Ollama quality gate locally:
.github/eval-ollama-config.jsonnpm run eval:ci:ollamaPrerequisites: Ollama installed, ollama serve running on 127.0.0.1:11434, and nomic-embed-text pulled.
Examples:
# Run against small golden set
npm run eval -- --dataset benchmarks/golden/small.json
# Compare against baseline
npm run eval:compare -- --against benchmarks/baselines/eval-baseline-summary.json --dataset benchmarks/golden/medium.json
# Sweep retrieval parameters
npm run eval -- --dataset benchmarks/golden/small.json --sweepFusionStrategy rrf,weighted --sweepHybridWeight 0.3,0.5,0.7 --sweepRrfK 30,60 --sweepRerankTopN 10,20
wrong-file, wrong-symbol, docs-tests-outranking-source, no-relevant-hit-top-k)Each run writes:
benchmarks/results/<timestamp>/
summary.jsonsummary.mdper-query.jsoncompare.json (when baseline/sweep used)benchmarks/golden/small.jsonbenchmarks/golden/medium.jsonbenchmarks/golden/large.jsonbenchmarks/budgets/github-models.json for the default GitHub Models workflow pathbenchmarks/budgets/default.json for explicit external provider overrides with baseline comparisonFull docs: docs/evaluation.md
Recent representative runs (plugin vs ripgrep vs ast-grep) on two medium repos:
Methodology for the snapshot below:
axios + expressdefinition, keyword-heavy) with scoped denominators shown in run reports--no-reindex, default)| Metric | Plugin | ripgrep | ast-grep (5/10 queries) |
|---|---|---|---|
| Hit@5 | 50% | 5% | 100% |
| MRR@10 | 0.48 | 0.04 | 0.90 |
| nDCG@10 | 0.48 | 0.08 | 0.93 |
| Latency p50 (ms) | 17.5 | 36.9 | 66.6 |
| Latency p95 (ms) | 30.9 | 44.1 | 70.7 |
--reindex)| Metric | Plugin | ripgrep | ast-grep (5/10 queries) |
|---|---|---|---|
| Hit@5 | 50% | 5% | 100% |
| MRR@10 | 0.48 | 0.04 | 0.98 |
| nDCG@10 | 0.48 | 0.07 | 0.98 |
| Latency p50 (ms) | 17.1 | 35.9 | 69.1 |
| Latency p95 (ms) | 30.4 | 43.7 | 75.1 |
ast-grep metrics are computed on its compatible query subset only (definition + keyword-heavy, 5/10 queries per repo). Plugin and ripgrep are scored on all 10 queries.
Interpretation:
For reproducible setup and commands (including with/without reindex), see:
docs/benchmarking-cross-repo.mdThe plugin automatically detects available credentials in this order:
nomic-embed-text)You can also use Custom to connect any OpenAI-compatible embedding endpoint (llama.cpp, vLLM, text-embeddings-inference, LiteLLM, etc.).
Each provider has different rate limits. The plugin automatically adjusts concurrency and delays:
| Provider | Concurrency | Delay | Best For |
|---|---|---|---|
| GitHub Copilot | 1 | 4s | Small codebases (<1k files) |
| OpenAI | 3 | 500ms | Medium codebases |
| 5 | 200ms | Medium-large codebases | |
| Ollama | 5 | None | Large codebases (10k+ files) |
| Custom | 3 | 1s | Any OpenAI-compatible endpoint |
For large codebases, use Ollama locally to avoid rate limits:
# Install the embedding model
ollama pull nomic-embed-text
// .opencode/codebase-index.json
{
"embeddingProvider": "ollama"
}
The built-in ollama provider uses Ollama's native /api/embeddings endpoint and is the simplest setup when you want to use nomic-embed-text.
For the built-in Ollama path, the plugin budgets nomic-embed-text against an observed effective input limit of about 2048 tokens, not the model's higher advertised theoretical context. This keeps batching and chunk text generation aligned with real Ollama embedding runtime behavior.
If you want to use a different Ollama embedding model through its OpenAI-compatible API, use the custom provider instead and set customProvider.baseUrl to http://127.0.0.1:11434/v1 so the plugin calls .../v1/embeddings.
The plugin is built for speed with a Rust native module (tree-sitter, usearch, SQLite). In practice, indexing and retrieval remain fast enough for interactive use on medium/large repositories.
For reproducible measurements on your machine, run: npx tsx benchmarks/run.ts.
Quick recommendation:
| Provider | Speed | Cost | Privacy | Best For |
|---|---|---|---|---|
| Ollama | Fastest | Free | Full | Large codebases, privacy-sensitive |
| GitHub Copilot | Slow (rate limited) | Free* | Cloud | Small codebases, existing subscribers |
| OpenAI | Medium | ~$0.0001/1K tokens | Cloud | General use |
| Fast | Free tier available | Cloud | Medium-large codebases | |
| Custom | Varies | Varies | Varies | Self-hosted or third-party endpoints |
*Requires active Copilot subscription
Set the provider in .opencode/codebase-index.json:
{ "embeddingProvider": "ollama" }
Credentials (if required) are read from environment variables (for example OPENAI_API_KEY or GOOGLE_API_KEY).
Custom (OpenAI-compatible)
Works with any server that implements the OpenAI /v1/embeddings API format (llama.cpp, vLLM, text-embeddings-inference, LiteLLM, etc.).
{
"embeddingProvider": "custom",
"customProvider": {
"baseUrl": "{env:EMBED_BASE_URL}",
"model": "nomic-embed-text",
"dimensions": 768,
"apiKey": "{env:EMBED_API_KEY}",
"maxTokens": 8192,
"timeoutMs": 30000,
"maxBatchSize": 64
}
}
Required fields: baseUrl, model, dimensions (positive integer). Optional: apiKey, maxTokens, timeoutMs (default: 30000), maxBatchSize (or max_batch_size) to cap inputs per /embeddings request for servers like text-embeddings-inference. {env:VAR_NAME} placeholders are resolved before config validation for fields that are actually used and throw if the referenced environment variable is missing or malformed.
Custom Ollama models via OpenAI-compatible API
If you are running Ollama locally and want to use an embedding model other than the built-in ollama setup, point the custom provider at Ollama's OpenAI-compatible base URL with the /v1 suffix:
{
"embeddingProvider": "custom",
"customProvider": {
"baseUrl": "http://127.0.0.1:11434/v1",
"model": "qwen3-embedding:0.6b",
"dimensions": 1024,
"apiKey": "ollama"
}
}
Notes:
/embeddings, so baseUrl should be http://127.0.0.1:11434/v1, not just http://127.0.0.1:11434."ollama" is fine.dimensions matches the actual output size of the model you pulled locally.Be aware of these characteristics:
| Aspect | Reality |
|---|---|
| Search latency | ~800-1000ms per query (embedding API call) |
| First index | Takes time depending on codebase size (e.g., ~30s for 500 chunks) |
| Requires API | Needs an embedding provider (Copilot, OpenAI, Google, or local Ollama) |
| Token costs | Uses embedding tokens (free with Copilot, minimal with others) |
| Best for | Discovery and exploration, not exhaustive matching |
Build:
npm run build
Register in Test Project (use file:// URL in opencode.json):
{
"plugin": [
"file:///path/to/opencode-codebase-index"
]
}
This loads directly from your source directory, so changes take effect after rebuilding.
For contribution workflow, standards, and release-label requirements, see CONTRIBUTING.md.
If you want to add support for a new language, see docs/adding-language-support.md for the full Rust + TypeScript checklist.
Quick path:
npm run build && npm run typecheck && npm run lint && npm run test:runTo ensure release notes reflect all merged work, this repo uses a draft-release workflow.
feature, bug, performance, documentation, dependencies, refactor, test, choresemver:major, semver:minor, or semver:patchRelease Label Check) and fail if no release category label is presentmain.git log --oneline vX.Y.Z..HEAD (or the previous release tag range) against the draft release notes so the release summary covers the full shipped delta, not just the current CHANGELOG.md Unreleased sectionCHANGELOG.mdpackage.json versionnpm run build && npm run typecheck && npm run lint && npm run test:rungh release create after reviewing draft content).PRs labeled skip-changelog are intentionally excluded from release notes.
├── src/
│ ├── index.ts # Plugin entry point
│ ├── mcp-server.ts # MCP server (Cursor, Claude Code, Windsurf)
│ ├── cli.ts # CLI entry for MCP stdio transport
│ ├── config/ # Configuration schema
│ ├── embeddings/ # Provider detection and API calls
│ ├── indexer/ # Core indexing logic + inverted index
│ ├── git/ # Git utilities (branch detection)
│ ├── tools/ # OpenCode tool definitions
│ ├── utils/ # File collection, cost estimation
│ ├── native/ # Rust native module wrapper
│ └── watcher/ # File/git change watcher
├── native/
│ └── src/ # Rust: tree-sitter, usearch, xxhash, SQLite
├── tests/ # Unit tests (vitest)
├── commands/ # Slash command definitions
├── skill/ # Agent skill guidance
└── .github/workflows/ # CI/CD (test, build, publish)
The Rust native module handles performance-critical operations:
Rebuild with: npm run build:native (requires Rust toolchain)
Pre-built native binaries are published for:
| Platform | Architecture | SIMD Acceleration |
|---|---|---|
| macOS | x86_64 | ✅ simsimd |
| macOS | ARM64 (Apple Silicon) | ✅ simsimd |
| Linux | x86_64 (GNU) | ✅ simsimd |
| Linux | ARM64 (GNU) | ✅ simsimd |
| Windows | x86_64 (MSVC) | ❌ scalar fallback |
Windows builds use scalar distance functions instead of SIMD — functionally identical, marginally slower for very large indexes. This is due to MSVC lacking support for certain AVX-512 intrinsics used by simsimd.
MIT
FAQs
Host-neutral semantic codebase search with embeddings, symbol discovery, and call-graph tooling
The npm package opencode-codebase-index receives a total of 931 weekly downloads. As such, opencode-codebase-index popularity was classified as not popular.
We found that opencode-codebase-index demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Security News
GPT-6 Astra hits 100% on ExploitBench and finds zero-days autonomously, while independent tests reveal scope violations and monitoring gaps.

Product
Socket can now send alerts and supply chain attack notifications to Microsoft Teams, with filters that route the right updates to each channel.

Security News
pnpm 12 rewrites the package manager in Rust, cutting install times by up to 90% while preserving pnpm 11 workflows and lockfiles.