
Company News
AWS Security Hub Adds Socket for Supply Chain Security
Socket is now in the AWS Security Hub Extended plan. Adopt it through AWS, apply committed spend, and block malicious open source packages.
vectora-agent
Advanced tools
Vectora is an open-source AI assistant (Apache 2.0) built for developers — local-first, self-hosted, and designed to run as a powerful sub-agent inside any MCP-compatible orchestrator (Claude Code, Claude Desktop, Paperclip, VS Code extensions).
At its core, Vectora solves the knowledge gap problem: LLMs don't know your codebase, your docs, or the latest versions of your stack. Vectora bridges that gap with RAG (Retrieval-Augmented Generation) — ingest your docs once, and every AI interaction becomes contextually aware.
web_cache collection and pass a curation gate (Cohere reranker + LLM judge) before being embedded — your curated knowledge base is never contaminated by unreviewed web results.Every message enters through a single entry point and is routed by the Orchestrator to the right specialized agent:
START
└─► orchestrator (responds inline OR delegates with task_query)
├─► [respond] → END
├─► [search] → search → search_tools → process_retrieval ↻ → END
├─► [coder] → coder → coder_tools ↻ → END
└─► [rag_subgraph] → rag_subgraph → orchestrator (synthesis) → END
| Agent | Responsibility | Tools |
|---|---|---|
| orchestrator | Primary LLM agent — responds directly OR delegates with an explicit task description | create_artifact, save_memory, get_memory, delete_memory |
| search | Web research, real-time info, builds knowledge base via cascading embeddings | web_search, fetch_url, vector_search |
| coder | File operations, terminal commands, code generation | file_read, file_edit, file_write, grep, list_dir, terminal |
| rag | Retrieval pipeline — retrieve → score → rerank/websearch → inject → orchestrator | vector_search, embedding, ingest_docs, manage_retriever (via subgraph) |
When the orchestrator routes to rag, a dedicated subgraph runs the full retrieval pipeline before synthesis:
rag_retrieve (vector_search)
└─► rag_decide (score threshold)
├─► rag_inject (score ≥ 0.7 — high confidence, inject directly)
├─► rag_rerank (score 0.4–0.7 — rerank with Cohere before inject)
└─► rag_websearch (score < 0.4 — fall back to web + auto-embed results)
Results are injected as a SystemMessage into context. The Orchestrator then synthesizes the final answer inline, without a separate agent hop.
Agents explicitly call create_artifact to persist structured documents (plans, specs, guides, architecture decisions) to ~/.vectora/artifacts/{session_id}/ as Markdown files. The tool returns structured metadata (path, title, type, session_id, timestamp) that the Orchestrator can reference in future turns.
After any web_search or fetch_url call, process_retrieval routes results through a curation gate before embedding:
web_persist_min_score are discarded.keep/discard verdict per document.Approved content is embedded into a dedicated web_cache collection, isolated from articles (user-curated content). The /rag panel shows the breakdown per collection, and manage_retriever lets you audit or remove cached web content at any time.
Vectora uses Cohere for embeddings (embed-multilingual-v3.0) and reranking (rerank-multilingual-v3.0). It offers a generous free tier with first-class LangChain integration.
Get your key: https://dashboard.cohere.com/api-keys
Vectora uses Tavily for real-time web search and URL content extraction. It offers a generous free tier optimized for AI agents.
Get your key: https://app.tavily.com/
| Provider | Free Tier | Get Key |
|---|---|---|
| Google Gemini ✅ Recommended | Yes | aistudio.google.com |
| Cohere | Yes | dashboard.cohere.com |
| Ollama (local) | No cost | ollama.ai |
| OpenAI | Paid | platform.openai.com |
| Anthropic | Paid | console.anthropic.com |
Install Vectora globally with uv:
uv tool install vectora-agent
On first run, the setup wizard will ask for your API keys and write them to ~/.vectora/.env.
vectora # starts chat (wizard runs automatically if no keys found)
To connect Vectora as an MCP sub-agent for Claude Code or Claude Desktop, add to your .mcp.json:
{
"mcpServers": {
"Vectora": {
"command": "vectora",
"args": ["mcp-server"]
}
}
}
Use this when you want Vectora running on a server and accessible from multiple machines or orchestrators via SSE.
Local (no domain):
cp .env.example .env
# Edit .env with your API keys
docker compose up -d
# SSE endpoint: http://localhost:8000/sse
VPS with Traefik (HTTPS + domain):
cp .env.example .env
# Edit .env with your API keys, VECTORA_DOMAIN and ACME_EMAIL
# Create the shared Traefik network if it doesn't exist yet
docker network create traefik-public
docker compose -f docker-compose.yml -f docker-compose.traefik.yml up -d
# SSE endpoint: https://vectora.yourdomain.com/sse
To connect from Claude Code or any MCP-compatible orchestrator:
{
"mcpServers": {
"Vectora": {
"url": "https://vectora.yourdomain.com/sse"
}
}
}
git clone https://github.com/brunosrz/vectora.git
cd vectora
uv sync
cp .env.example .env
# Edit .env with your API keys
uv run vectora
vectora [options] Start chat (resume last session for this directory)
vectora mcp-server Start MCP server (stdio)
vectora traces View observability traces
vectora sessions List all saved sessions
vectora config Show current configuration
vectora config --set KEY=VALUE Edit a setting
Options:
--model MODEL Switch LLM model (provider auto-detected). Persists.
--ollama Force Ollama provider (for arbitrary local model names)
--session ID Resume a specific session by 6-digit ID
--new Force a new session
--verbosity N Verbosity level 0–5 (0=silent, 5=debug panel). Persists.
--version Show version
| Command | Description |
|---|---|
/help | Show quick help |
/list | Show all commands |
/tools | List available tools |
/model | List or switch models |
/debug [0-5] | Set verbosity level (tool calls, routing decisions, log panel) |
/new | Start a new session |
/sessions | List all sessions |
/session <id> | Switch to a specific session |
/quit | Exit |
Input shortcuts: Enter sends, Alt+Enter or Shift+Enter adds a line break.
16 tools across 5 categories, always available to all agents:
| Category | Tools | Primary Agent |
|---|---|---|
| Web | web_search, fetch_url | search |
| RAG | vector_search, embedding, ingest_docs, manage_retriever | search / RAG subgraph |
| Files | file_read, file_edit, file_write, grep, list_dir, terminal | coder |
| Artifacts | create_artifact | orchestrator |
| Memory | save_memory, get_memory, delete_memory | orchestrator / coder |
All data is stored locally in ~/.vectora/:
~/.vectora/
├── .env # API keys (secrets — never commit)
├── settings.json # Runtime preferences (provider, model, verbosity)
├── data/
│ ├── vectora.db # Sessions, memories, LangGraph checkpoints (SQLite)
│ ├── embedding_queue.db # Async embedding queue (SQLite)
│ ├── traces.db # Internal observability spans (SQLite)
│ └── lancedb/ # Vector store for RAG (LanceDB)
├── artifacts/ # Auto-detected plans, specs, guides
│ └── {session_id}/
│ └── *.md
├── keys/ # Reserved for future key management
└── logs/
├── vectora.jsonl # Structured JSON logs
└── session_*.md # Exported session audit trails
Separation of concerns:
~/.vectora/.env — secrets (API keys). Never versioned.~/.vectora/settings.json — non-secret runtime preferences (active provider, model, verbosity, last session per directory). Managed by vectora config.| Layer | Technology |
|---|---|
| Language | Python 3.14+ managed by uv |
| Agent Framework | LangChain + LangGraph |
| Agent Pattern | Orchestrator + Specialized Workers (search / coder) + RAG Subgraph |
| Vector Store | LanceDB — file-based, zero-config |
| Embeddings | Cohere — embed-multilingual-v3.0 + rerank-multilingual-v3.0 |
| Persistence | SQLite via aiosqlite + LangGraph Checkpointer |
| Context Protocol | MCP via FastMCP |
| Terminal UI | Rich + prompt-toolkit |
| Observability | LangSmith (optional) |
API keys go in ~/.vectora/.env (created by the setup wizard) or a project-local .env:
# LLM Provider (auto-detected from available keys if not set)
LLM_PROVIDER=google-genai
GOOGLE_API_KEY=your_key_here
# Required: RAG embeddings + reranking
COHERE_API_KEY=your_key_here
# Required: Web search + URL extraction
TAVILY_API_KEY=your_key_here
# Optional: Tracing
LANGSMITH_TRACING=false
LANGSMITH_API_KEY=your_key_here
LANGSMITH_PROJECT=vectora
Runtime preferences (model, verbosity, session history) are managed in ~/.vectora/settings.json via vectora config or the /model and /debug chat commands — no need to touch .env for these.
Apache 2.0. See LICENSE.
FAQs
Vectora - Advanced AI Assistant with RAG and MCP capabilities
We found that vectora-agent demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.
Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Company News
Socket is now in the AWS Security Hub Extended plan. Adopt it through AWS, apply committed spend, and block malicious open source packages.

Research
/Security News
Popular npm packages keyv and cacheable compromised.

Security News
A misconfiguration gave three Anthropic models internet access, and one, believing it was in a simulation, shipped a credential-stealing package to PyPI.