New:Microsoft Teams Notifications Are Now Available in Socket.Learn more
Get Started

@agentskit/rag

Package Overview
Dependencies
Maintainers
1
Versions
36
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

@agentskit/rag

Plug-and-play retrieval-augmented generation for AgentsKit.

latest
Source
npmnpm
Version
0.5.6
Version published
Weekly downloads
167
-35.02%
Maintainers
1
Weekly downloads
 
Created
Source

@agentskit/rag

AgentsKit

Plug-and-play retrieval-augmented generation: chunk documents, embed them, and retrieve the right context at query time.

npm version npm downloads bundle size license stability GitHub stars

Tags: ai · agents · llm · agentskit · rag · retrieval · vector-search · embeddings · ai-agents · semantic-search · knowledge-base

Verified proof

How this fits the ecosystem

@agentskit/rag is the retrieval layer: load documents, chunk them, embed them, rerank results, and feed precise context back to agents.

  • AgentsKit: compose it with the other packages in this repo to build agents from small, swappable parts.
  • Registry: look for ready agents and templates that already use this layer at registry.agentskit.io.
  • Playbook: learn the production patterns behind this layer at playbook.agentskit.io.
  • AKOS: run the same concepts with enterprise deployment, governance, and observability at akos.agentskit.io.

Docs: package guide · agent handoff

Why rag

  • Your data, your agent — no fine-tuning required; ingest plain text and query with natural language
  • Composable stack — uses any EmbedFn and any VectorMemory from @agentskit/adapters and @agentskit/memory; swap either layer without touching RAG logic
  • Retriever-readycreateRAG() returns a Retriever you pass to @agentskit/runtime or useChat so context is injected automatically
  • Tune chunking without a PhDchunkSize, chunkOverlap, or a custom split function — three knobs that cover 95% of use cases

Install

npm install @agentskit/rag @agentskit/memory @agentskit/adapters

The file-backed example also needs the optional vectra peer. Add vectra to the install command when using fileVectorMemory; the runtime integration example additionally needs @agentskit/runtime.

Quick example

import { createRAG } from '@agentskit/rag'
import { openaiEmbedder } from '@agentskit/adapters'
import { fileVectorMemory } from '@agentskit/memory'

const rag = createRAG({
  embed: openaiEmbedder({ apiKey: process.env.OPENAI_API_KEY! }),
  store: fileVectorMemory({ path: './vectors' }),
})

await rag.ingest([
  { id: 'doc-1', content: 'AgentsKit is a JavaScript agent toolkit...' },
])

const docs = await rag.search('How does AgentsKit work?', { topK: 5 })
console.log(docs)

With runtime (retriever)

Pass the RAG instance as retriever so the runtime injects retrieved context into the task:

import { createRuntime } from '@agentskit/runtime'
import { openai } from '@agentskit/adapters'

const runtime = createRuntime({
  adapter: openai({ apiKey: process.env.OPENAI_API_KEY!, model: 'gpt-4o' }),
  retriever: rag,
})

const result = await runtime.run('Explain the AgentsKit architecture based on ingested docs')
console.log(result.content)

You can also call rag.retrieve({ query, messages }) to satisfy the core Retriever contract (for example from a custom controller).

Features

  • createRAG({ embed, store }) — single entry point for ingest + retrieve.
  • rag.ingest(docs) — chunk, embed, and store documents.
  • rag.search(query, { topK }) — semantic similarity search.
  • rag.retrieve({ query, messages })Retriever contract v1 for runtime/controller injection.
  • Configurable chunking: chunkSize, chunkOverlap, custom split.
  • Works with any EmbedFn and any VectorMemory.
  • Rerankers: createRerankedRetriever (Voyage, Jina, custom RerankFn, BM25 default), createHybridRetriever (vector + BM25 blend), standalone bm25Score. Recipe.
  • Document loaders: loadUrl, loadGitHubFile, loadGitHubTree, loadNotionPage, loadConfluencePage, loadGoogleDriveFile, loadPdf, loadS3, loadGcs, loadDropbox, and loadOneDrive. Recipe.
  • Loader resilience: HTTP/network and response-body read/parse failures surface as RagError (AK_RAG_LOAD_FAILED). Every remote request and body read has a finite timeout and byte limit by default; both are configurable through timeoutMs and maxResponseBytes. Optional signal aborts are never swallowed as a per-object skip. Tree/list loaders may return partial success when at least one eligible download succeeded; if every attempted eligible download failed, they throw. Missing/invalid S3 object bodies count as failed downloads. Pagination that reports more data without a new cursor/token throws (no silent truncation). loadNotionPage follows Notion has_more / next_cursor with start_cursor until complete (preserving block order; incomplete or repeated cursors throw). loadUrl requires an HTTPS origin in allowedOrigins; it does not follow redirects. Non-positive / non-finite maxFiles yields [].
  • Score contracts: scoreless search/rerank results keep order. When any score is present, every result must have a finite numeric score and is sorted descending — mixed or non-finite scores throw (never fabricate -Infinity). Malformed Voyage/Jina/custom reranker output throws AK_RAG_RERANK_FAILED. Optional signal on voyageReranker / jinaReranker is forwarded to fetch; request/body aborts remain AK_RAG_RERANK_FAILED. bm25Score sanitizes invalid k1/b to documented defaults and always emits finite scores. Hybrid relative weights are normalized to a finite pair that sums to 1 (both zero → 0.5/0.5).
  • Chunk/config safety: invalid chunkSize / chunkOverlap / topK values are sanitized so chunking always terminates and search never sends non-finite limits to the store.
  • Ingestion behavior: rag.ingest embeds chunks serially and sends one vector-store batch per call. For large corpora, batch documents in the caller and persist progress between calls; the package does not silently add concurrency or an unbounded background queue.

S3 in Expo and React Native runtimes

Node consumers may install @aws-sdk/client-s3 and let loadS3 resolve it lazily. Browser, Expo/Metro, and React Native bundles keep that peer out of the universal entry; pass the command constructors explicitly when invoking the loader:

import { GetObjectCommand, ListObjectsV2Command, S3Client } from '@aws-sdk/client-s3'
import { loadS3 } from '@agentskit/rag'

await loadS3({
  client: new S3Client({}),
  bucket: 'knowledge',
  commands: { GetObjectCommand, ListObjectsV2Command },
})

Ecosystem

PackageRole
@agentskit/coreRetriever, VectorMemory, types
@agentskit/memoryVector backends (fileVectorMemory, etc.)
@agentskit/adaptersopenaiEmbedder and other embedders
@agentskit/runtimeretriever integration for agents
@agentskit/reactuseChat + chat UI with the same core types

Contributors

AgentsKit contributors

License

MIT — see LICENSE.

Docs

Full documentation · GitHub

Maturity and compatibility

  • Stability: beta — see docs/STABILITY.md
  • Node.js 20+ and TypeScript strict mode
  • Published as @agentskit/rag

Contributing

See CONTRIBUTING.md and the monorepo LICENSE.

Keywords

agentskit

FAQs

Package last updated on 03 Sep 2026

Related posts