
Company News
AWS Security Hub Adds Socket for Supply Chain Security
Socket is now in the AWS Security Hub Extended plan. Adopt it through AWS, apply committed spend, and block malicious open source packages.
docsonar
Advanced tools
Chat with your folders. A local document search MCP server — your files never leave your machine.
docsonar indexes folders of documents into a single SQLite database and gives any MCP client (Claude Desktop, Claude Code, and others) hybrid keyword + semantic search over them. Fully offline, zero configuration, read-only by design.
Add to your project's .mcp.json (or ~/.claude.json for all projects):
{
"mcpServers": {
"docsonar": {
"command": "uvx",
"args": ["docsonar"]
}
}
}
Add to claude_desktop_config.json (Settings → Developer → Edit Config):
{
"mcpServers": {
"docsonar": {
"command": "uvx",
"args": ["docsonar"]
}
}
}
Running from a source checkout instead: "command": "uv", "args": ["run", "--directory", "/path/to/docsonar", "docsonar"].
Then just talk to it: "Add my ~/Documents/notes folder and find everything about quarterly planning." The first index downloads the embedding model (~130 MB, one time); keyword search works immediately while that happens.
For HTTP instead of stdio: uvx docsonar --transport http --port 8365.
| Tool | Purpose |
|---|---|
add_folder(path, include_globs?, exclude_globs?) | Register a folder; indexing starts in the background |
remove_folder(path) | Unregister and purge its index data |
list_folders() | Registered folders with counts and last index time |
search(query, top_k?, folder?, file_type?, mode?) | Hybrid keyword+semantic search (RRF-fused); ranked passages with file path, heading/page location, and snippet |
read_file(path, start_line?, end_line?) | File text (extracted text for pdf/docx/html); refuses paths outside registered folders |
find_similar(path, top_k?) | Documents most similar in meaning to a given file |
reindex(folder?, force?) | Incremental refresh: new/changed files reindexed, deleted files purged |
index_status() | Index totals, embedding model state, background job progress, failed files |
Supported formats: txt, md, pdf (page-aware, cites p. 12), docx, html.
flowchart LR
Client["MCP client<br/>(Claude Desktop / Code)"] <-->|stdio or HTTP| Server["FastMCP server<br/>8 tools"]
Server --> Sec["path security<br/>(registered folders only)"]
Server --> Search["hybrid search<br/>BM25 + cosine KNN + RRF"]
Server --> Worker["background indexer<br/>(worker thread + queue)"]
Worker --> Parsers["parsers<br/>txt · md · pdf · docx · html"]
Parsers --> Chunker["heading/page-aware chunker<br/>~400 tokens, 60 overlap"]
Chunker --> DB[("SQLite<br/>FTS5 + sqlite-vec")]
Search --> DB
Embed["sentence-transformers<br/>bge-small-en-v1.5 (local)"] --> Worker
Embed --> Search
Every chunk stores its location (heading path like Setup > Windows, or PDF page range), so search results can cite report.pdf, p. 12.
mode lets the caller force either side.search returns enough per hit (path, location, score, snippet) to decide what to read next without another round trip; results carry stable chunk_ids; every degradation is reported in-band via note/embedding_note fields instead of failing. There's no "answer the question" tool on purpose — the calling model does the reasoning; docsonar does retrieval.Synthetic corpus of 200 markdown files (600 chunks); Intel Core Ultra 5 125U (laptop, CPU-only), Windows 11, Python 3.12. Reproduce with uv run python scripts/benchmark.py.
| Operation | Result |
|---|---|
| Index, keyword-only | 0.7 s (≈280 files/s) |
| Index, with embeddings | 59 s (≈3.4 files/s — embedding-bound) |
| Incremental reindex, nothing changed | 0.03 s |
| Search, keyword | 8.9 ms median |
| Search, semantic | 35 ms median |
| Search, hybrid | 54 ms median |
| Embedding model load (once per process) | ~21 s |
| Database size | 2.6 MB (1.1 MB keyword-only) |
Zero config needed. To customize, create config.toml in the platform config directory (Windows: %LOCALAPPDATA%\docsonar\, macOS: ~/Library/Application Support/docsonar/, Linux: ~/.config/docsonar/):
embedding_model = "BAAI/bge-small-en-v1.5" # any sentence-transformers model
chunk_target_tokens = 400
chunk_max_tokens = 512
chunk_min_tokens = 100
chunk_overlap_tokens = 60
max_file_size_mb = 50
extra_ignore_dirs = ["Archive"]
CLI flags: --transport stdio|http, --host, --port, --db-path, --config.
The index database lives in the platform data directory (Windows: %LOCALAPPDATA%\docsonar\, macOS: ~/Library/Application Support/docsonar/, Linux: ~/.local/share/docsonar/).
Requires uv (it installs the right Python automatically):
# Windows
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
git clone https://github.com/ribhav-jain/docsonar.git
cd docsonar
uv sync --dev # creates .venv and installs everything, incl. dev tools
uv run pytest # full suite (~2 s, no model download needed)
uv run pytest -v # verbose, one line per test
uv run pytest tests/test_search.py # one file
uv run pytest -k incremental # tests matching a keyword
uv run ruff check . # lint
uv run mypy src tests # strict typecheck
uv run python scripts/benchmark.py # perf numbers (downloads the real model)
uv run docsonar # stdio — for MCP clients (Claude Desktop/Code)
uv run docsonar --transport http --port 8365 # HTTP — for manual testing (Postman, curl)
VS Code: press F5 — launch configs for the HTTP server and the test suite are in .vscode/launch.json.
With the HTTP server running, the endpoint is http://127.0.0.1:8365/mcp. Recent Postman versions can connect directly (New → MCP Request, transport HTTP) and show all 8 tools as forms. For raw HTTP, MCP is JSON-RPC over POST with two setup calls, then tool calls. Send every request with headers Content-Type: application/json and Accept: application/json, text/event-stream:
// 1. initialize — copy the `mcp-session-id` RESPONSE header and send it
// back as a request header on every call below
{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-03-26","capabilities":{},"clientInfo":{"name":"postman","version":"1.0"}}}
// 2. initialized notification (expect 202, empty body)
{"jsonrpc":"2.0","method":"notifications/initialized"}
// 3. list the tools
{"jsonrpc":"2.0","id":2,"method":"tools/list"}
Then call tools — the repo ships a sample corpus in examples/sample-docs to play with (use forward slashes in JSON to avoid escaping; they work on Windows too):
// Register the sample folder (first ever call downloads the embedding model, ~30 s)
{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"add_folder",
"arguments":{"path":"C:/path/to/docsonar/examples/sample-docs"}}}
// Watch indexing progress until indexing.active is false
{"jsonrpc":"2.0","id":4,"method":"tools/call","params":{"name":"index_status","arguments":{}}}
// Hybrid search — finds the expense-policy PDF, cites its page
{"jsonrpc":"2.0","id":5,"method":"tools/call","params":{"name":"search",
"arguments":{"query":"meal reimbursement limit","top_k":3,"mode":"hybrid"}}}
// Semantic search — no keyword overlap needed
{"jsonrpc":"2.0","id":6,"method":"tools/call","params":{"name":"search",
"arguments":{"query":"how hot should water be for green tea","mode":"semantic"}}}
// Read a hit (works on PDFs too — returns extracted text)
{"jsonrpc":"2.0","id":7,"method":"tools/call","params":{"name":"read_file",
"arguments":{"path":"C:/path/to/docsonar/examples/sample-docs/expense-policy.pdf"}}}
// "More like this"
{"jsonrpc":"2.0","id":8,"method":"tools/call","params":{"name":"find_similar",
"arguments":{"path":"C:/path/to/docsonar/examples/sample-docs/espresso-guide.md","top_k":3}}}
// Refresh after files change on disk (force:true rebuilds everything)
{"jsonrpc":"2.0","id":9,"method":"tools/call","params":{"name":"reindex","arguments":{}}}
// Clean up — unregisters the folder and purges its index data
{"jsonrpc":"2.0","id":10,"method":"tools/call","params":{"name":"remove_folder",
"arguments":{"path":"C:/path/to/docsonar/examples/sample-docs"}}}
Security check worth trying: read_file with any path outside a registered folder (e.g. C:/Windows/System32/drivers/etc/hosts) returns a refusal, not file content.
CI runs lint, typecheck, and the test suite on Linux, macOS, and Windows.
MIT
FAQs
Local document search MCP server — chat with your folders, fully offline.
The pypi package docsonar receives a total of 28 weekly downloads. As such, docsonar popularity was classified as not popular.
We found that docsonar demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.
Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Company News
Socket is now in the AWS Security Hub Extended plan. Adopt it through AWS, apply committed spend, and block malicious open source packages.

Research
/Security News
Popular npm packages keyv and cacheable compromised.

Security News
A misconfiguration gave three Anthropic models internet access, and one, believing it was in a simulation, shipped a credential-stealing package to PyPI.