MemStack
The open-source memory layer for AI agents — store, retrieve, summarize, and prune.

npm install @memstack/core
The problem: AI agents forget. Every interaction starts from zero. You either stuff everything into the context window (expensive, slow, degrades output quality) or the agent has no memory of past conversations.
What MemStack does: A persistent memory pipeline that lives between your agent and the LLM. It stores every interaction, retrieves only what's relevant, summarizes old memories to save tokens, and prunes stale ones automatically. One method call, no infrastructure required.
Think of it as the open-source alternative to Mem0 — pluggable storage, bring your own LLM, zero vendor lock-in.
Table of Contents
Why MemStack
LLMs have context windows, not memory. The difference matters.
| Stuff everything in context | Cost is O(n²). 100 conversations = thousands of tokens = dollars per call. Quality degrades from "lost in the middle" effect. |
| Use a vector DB directly | You get similarity search. You don't get summarization, pruning, recency weighting, deduplication, or token budget management. You're building the pipeline yourself. |
| Use Mem0 | Proprietary, cloud-only with their hosted API. You don't control where your data lives. |
| Use MemStack | Full pipeline. Pluggable everything. Your data, your infrastructure. Open source. |
What MemStack handles that raw vector DBs don't:
- Summarization — compress 100 old interactions into one paragraph, keep meaning, save tokens
- Recency weighting — recent memories matter more; MemStack sorts them higher
- Importance scoring — not all memories are equal; high-importance ones survive pruning
- Deduplication — identical or near-identical memories are collapsed in context assembly
- Token budget —
compileContext() tells you how many tokens you're spending before the LLM call
- Memory-type routing — interactions, summaries, observations treated differently at retrieval time
- Auto-pruning — old, low-importance memories clean themselves up
Quick Start
import { MemStack, OpenAILLMAdapter, OpenAIEmbeddingAdapter, InMemoryStorageAdapter } from "@memstack/core";
const memstack = new MemStack({
llm: new OpenAILLMAdapter({ apiKey: process.env.OPENAI_API_KEY! }),
embedding: new OpenAIEmbeddingAdapter({ apiKey: process.env.OPENAI_API_KEY! }),
storage: new InMemoryStorageAdapter(),
});
await memstack.memory.store({
actorId: "support-bot-42",
content: "User reports login failing with error 503 on Chrome 125.",
tags: ["login", "bug", "chrome"],
importance: 0.8,
});
const memories = await memstack.memory.retrieve({
actorId: "support-bot-42",
query: "login error",
strategy: "hybrid",
});
const ctx = await memstack.memory.compileContext({
actorId: "support-bot-42",
maxTokens: 2000,
});
const llmResponse = await llm.complete({
system: `You are a support bot. Here is what you remember:\n${ctx.systemPrompt}`,
user: "The user is back and still can't log in. What do you do?",
});
The Memory Pipeline
MemStack's core is a five-stage pipeline. Each stage can be used independently.
1. Store
Every agent interaction becomes a Memory with metadata that controls how it's retrieved, summarized, and pruned later.
interface Memory {
id: string;
actorId: string;
memoryType: MemoryType;
content: string;
importance: number;
emotionalValence: number;
tags: string[];
embedding?: number[];
metadata?: Record<string, unknown>;
expiresAt?: Date;
sourceId?: string;
createdAt: Date;
}
await ms.memory.store({
actorId: "agent-7",
content: "Customer asked about refund policy for Q2 purchases.",
tags: ["billing", "refund"],
});
await ms.memory.storeBatch([
{ actorId: "agent-7", content: "First interaction" },
{ actorId: "agent-7", content: "Second interaction" },
{ actorId: "agent-7", content: "Third interaction" },
]);
2. Retrieve
Pull back what's relevant — by keyword, by meaning (semantic), by recency, or by importance.
const memories = await ms.memory.retrieve({
actorId: "agent-7",
query: "refund policy",
strategy: "hybrid",
limit: 10,
memoryTypes: ["interaction"],
tags: ["billing"],
});
Strategy behavior:
recent | Newest first | No | Knowing what just happened |
important | Highest importance first | No | Filtering noise, keeping signal |
semantic | Cosine similarity to query | Yes | "Find memories about X" |
hybrid | Semantic + importance blend | Yes | Best of both worlds |
No embedding adapter? semantic and hybrid fall back to keyword matching + importance sort. No API costs, just less precise.
3. Compile Context
The killer feature. compileContext() takes the retrieval results and assembles an LLM-ready system prompt — deduplicated, sorted by recency and importance, with a token estimate so you know the cost before calling the LLM.
const ctx = await ms.memory.compileContext({
actorId: "agent-7",
maxTokens: 2000,
memoryTypes: ["interaction", "summary"],
});
console.log(ctx.tokenEstimate);
const response = await llm.complete({
system: ctx.systemPrompt,
user: userMessage,
});
compileContext() is the difference between "we have a vector DB" and "we have agent memory." It handles deduplication, token budgeting, and the recent-vs-important split that makes context useful.
4. Summarize
When an actor has hundreds of interactions, retrieval gets expensive and context gets bloated. Summarization compresses old interactions into a single paragraph using the configured LLM.
const { summary, deletedCount } = await ms.memory.summarize({
actorId: "agent-7",
olderThan: new Date(Date.now() - 7 * 86400000),
skipMostRecent: 10,
targetCount: 50,
memoryTypes: ["interaction"],
keepOriginals: false,
});
console.log(deletedCount);
Auto-summarization: Set summarizationThreshold in config (default: 100). Every 100th interaction for an actor triggers summarization automatically.
Warning: keepOriginals: false deletes the summarized memories. Set keepOriginals: true to preserve them alongside the summary.
Custom summarization prompt:
const ms = new MemStack({
llm,
defaults: {
summarizationPrompt:
"You are an enterprise support memory compressor. Highlight: customer name,
product, severity, resolution status, and any open issues.",
},
});
5. Prune
Not all memories deserve to live forever. Pruning removes low-value memories to keep storage and retrieval fast.
await ms.memory.prune({ type: "byAge", maxAge: 30 * 86400000 });
await ms.memory.prune({ type: "byImportance", minImportance: 0.3 });
await ms.memory.prune({ type: "byCount", maxPerActor: 500 });
await ms.memory.prune({ type: "byType", memoryTypes: ["observation"] });
await ms.memory.prune({
type: "custom",
shouldRemove: (memory) => memory.content.includes("[RESOLVED]"),
});
const { wouldPrune, count } = await ms.memory.dryRunPrune({
type: "byAge",
maxAge: 86400000,
});
console.log(`Would remove ${count} memories:`, wouldPrune);
Auto-prune on every process() call by setting pruneStrategy in config:
const ms = new MemStack({
llm,
defaults: {
pruneStrategy: { type: "byImportance", minImportance: 0.05 },
},
});
Real-World Use Cases
Support Agent
async function handleMessage(customerId: string, message: string) {
await ms.memory.store({
actorId: `customer:${customerId}`,
content: message,
importance: detectUrgency(message),
tags: classifyIntent(message),
});
const ctx = await ms.memory.compileContext({
actorId: `customer:${customerId}`,
maxTokens: 1500,
});
const response = await llm.complete({
system: `You are a support agent. Customer history:\n${ctx.systemPrompt}`,
user: message,
});
return response.text;
}
RAG Pipeline
for (const doc of documents) {
await ms.memory.store({
actorId: "knowledge-base",
content: doc.text,
memoryType: "observation",
metadata: { source: doc.url, section: doc.section },
});
}
const relevantDocs = await ms.memory.retrieve({
actorId: "knowledge-base",
query: "How does authentication work?",
strategy: "semantic",
limit: 5,
});
const ctx = await ms.memory.compileContext({
actorId: "knowledge-base",
memoryTypes: ["observation"],
});
const answer = await llm.complete({
system: `Answer using only these documents:\n${ctx.systemPrompt}`,
user: "How does authentication work?",
});
Multi-User Chatbot
async function chat(userId: string, message: string) {
await ms.memory.store({
actorId: userId,
content: message,
});
const ctx = await ms.memory.compileContext({
actorId: userId,
maxTokens: 1000,
});
return llm.complete({
system: `You are a friendly assistant. Conversation history with this user:\n${ctx.systemPrompt}`,
user: message,
});
}
const total = await ms.memory.count();
const userCount = await ms.memory.count({ actorId: "user-42" });
Memory Type Reference
interaction | Default. Direct exchanges between agent and user/other agent. | "User asked about billing." |
summary | Compressed collection of old interactions. Created by summarize(). | "Over 3 weeks, user reported 5 login failures..." |
observation | Passive knowledge — facts, documents, things the agent knows but didn't interact with. | "Company refund policy is 30 days from purchase." |
fact | Verified knowledge — discrete truths the agent has confirmed. | "The user's subscription tier is Enterprise." |
reflection | Self-generated insight — the agent thinking about its own experiences. | "I tend to over-explain billing policies — should be more concise." |
Types control retrieval behavior — compileContext() treats interaction and summary differently from observation. Use types to separate "what happened" from "what I know."
Retrieval Strategies
Four strategies, each with a purpose:
await ms.memory.retrieve({ actorId: "x", strategy: "recent", limit: 3 });
await ms.memory.retrieve({ actorId: "x", strategy: "important" });
await ms.memory.retrieve({ actorId: "x", query: "login bug", strategy: "semantic" });
await ms.memory.retrieve({ actorId: "x", query: "login bug", strategy: "hybrid" });
Choosing a strategy:
- Use
recent for chatbots, ongoing conversations, anything time-sensitive
- Use
important for long-running agents where signal-to-noise matters
- Use
semantic for RAG, document search, knowledge base queries
- Use
hybrid for most agent memory — it balances meaning with significance
Embeddings
Embeddings power semantic search. They're optional — without them, retrieval uses keyword matching.
With embeddings (configure an EmbeddingProvider): each store() computes a vector. retrieve() with "semantic" or "hybrid" uses cosine similarity ranking.
Without embeddings: everything still works — retrieval falls back to importance + recency + keyword filters. No API costs, no setup.
Batch embedding: storeBatch() sends all texts in one embedding API call, reducing cost and latency.
const ms = new MemStack({
llm,
embedding: new OpenAIEmbeddingAdapter({ apiKey }),
defaults: { embedOnStore: false },
});
Adapters
MemStack is provider-agnostic. Every boundary is an interface — bring your own LLM, embedding model, and storage backend.
LLM Adapters
Used by summarize() and compileContext(). Ships with OpenAI, Anthropic, Ollama, and Groq built-in — and via baseURL, the OpenAI adapter works with any OpenAI-compatible API (DeepSeek, Mistral, Gemini, Together AI, Perplexity, Fireworks, xAI, and dozens more).
import { OpenAILLMAdapter } from "@memstack/core";
const llm = new OpenAILLMAdapter({ apiKey: "..." });
const deepseek = new OpenAILLMAdapter({ apiKey: "...", baseURL: "https://api.deepseek.com/v1" });
const mistral = new OpenAILLMAdapter({ apiKey: "...", baseURL: "https://api.mistral.ai/v1" });
const together = new OpenAILLMAdapter({ apiKey: "...", baseURL: "https://api.together.xyz/v1" });
import { AnthropicLLMAdapter } from "@memstack/core";
const llm = new AnthropicLLMAdapter({
apiKey: process.env.ANTHROPIC_API_KEY!,
defaultModel: "claude-sonnet-4-5-20250929",
});
import { OllamaLLMAdapter } from "@memstack/core";
const llm = new OllamaLLMAdapter({
baseURL: "http://localhost:11434",
defaultModel: "llama3.2",
});
Embedding Adapters
Used by semantic retrieval. Ships with OpenAI and Cohere built-in — and via baseURL, the OpenAI adapter works with any OpenAI-compatible embedding API (Together AI, Voyage AI, Jina, Nomic, and more).
import { OpenAIEmbeddingAdapter, CohereEmbeddingAdapter } from "@memstack/core";
new OpenAIEmbeddingAdapter({ apiKey: "...", model: "text-embedding-3-small" });
new CohereEmbeddingAdapter({ apiKey: "..." });
new OpenAIEmbeddingAdapter({ apiKey: "...", baseURL: "https://api.voyageai.com/v1", model: "voyage-3" });
Storage Adapters
MemStack ships with 11 production-ready storage adapters (7 experimental) — every major backend, zero peer dependencies, all client-injected.
Built-in (zero external deps):
InMemoryStorageAdapter | In-memory Map | Testing, prototyping |
DiskStorageAdapter | Local JSON files | Simple local persistence |
MarkdownStorageAdapter | Append-only .md files | Human-readable, git-diffable, debug-friendly |
HybridStorageAdapter | Compose any two StorageProviders | Cache + durable, edge + durable |
Relational / SQL:
PostgresStorageAdapter | PostgreSQL + pgvector | HNSW native |
SQLiteStorageAdapter | SQLite (better-sqlite3) | In-memory cosine |
TursoStorageAdapter | Turso (libsql) | DiskANN native |
Aggregators:
Mem0StorageAdapter | Mem0 OSS or Cloud |
ZepStorageAdapter | Zep Cloud or Community Edition |
Vector databases:
QdrantStorageAdapter | Qdrant |
PineconeStorageAdapter | Pinecone |
ChromaStorageAdapter | ChromaDB |
WeaviateStorageAdapter | Weaviate |
LanceDBStorageAdapter | LanceDB |
MongoDBStorageAdapter | MongoDB Atlas Vector Search |
Cache / KV:
RedisStorageAdapter | Redis (ioredis) |
UpstashStorageAdapter | Upstash Redis + Vector |
Graph:
Quick-start per backend:
import { PostgresStorageAdapter } from "@memstack/core";
const storage = new PostgresStorageAdapter({ connectionString: "postgres://..." });
import Database from "better-sqlite3";
import { SQLiteStorageAdapter } from "@memstack/core";
const storage = new SQLiteStorageAdapter({ db: new Database("memory.db") });
import Redis from "ioredis";
import { RedisStorageAdapter } from "@memstack/core";
const storage = new RedisStorageAdapter({ redis: new Redis() });
import { MarkdownStorageAdapter } from "@memstack/core";
const storage = new MarkdownStorageAdapter({ dir: "./memories" });
import { HybridStorageAdapter } from "@memstack/core";
const storage = new HybridStorageAdapter({
cache: new RedisStorageAdapter({ redis: new Redis() }),
durable: new PostgresStorageAdapter({ connectionString: "postgres://..." }),
});
Custom storage:
import type { StorageProvider, MemoryStoreInput } from "@memstack/core";
class MyStorage implements StorageProvider {
async store(input: MemoryStoreInput): Promise<Memory> { }
async get(id: string): Promise<Memory | null> { }
async retrieve(query: MemoryRetrieveQuery, embedding?: number[]): Promise<Memory[]> { }
async count(filter?: MemoryCountFilter): Promise<number> { }
async delete(id: string): Promise<void> { }
async deleteMany(ids: string[]): Promise<number> { }
async storeBatch(inputs: MemoryStoreInput[]): Promise<Memory[]> { }
async initialize(): Promise<void> { }
async close(): Promise<void> { }
}
Backend Comparison
| InMemory | Cosine in-memory | Yes | Testing, prototyping |
| Disk (JSON) | Keyword + importance | Yes | Simple local persistence |
| Markdown | Keyword + importance | No | Human-readable, git-diffable |
| Postgres | pgvector HNSW | Yes | Production relational |
| SQLite | Cosine in-memory | Yes | Local dev, solo apps |
| Turso | DiskANN native | Yes | Edge/serverless |
| Redis | RediSearch KNN (auto-detect) | Yes | Sub-5ms hot session state |
| Upstash | Native vector (vector mode) | No | CF Workers, Vercel Edge |
| Qdrant | ANN native | No | Best filtered search |
| Pinecone | ANN native | No | Zero-ops managed |
| Chroma | Native | No | LangChain prototyping |
| Weaviate | BM25 + vector hybrid | No | Hybrid search |
| LanceDB | DiskANN native | No | Embedded local vector |
| MongoDB | Atlas Vector Search | No | Existing MongoDB deployments |
| Neo4j | Neo4j vector index | No | Relationship-aware agents |
| Hybrid | Delegates to cache/durable | If durable supports | Read-through cache pattern |
| Mem0 | Delegates to Mem0 | No | Multi-backend via Mem0 |
| Zep | Graphiti temporal graph | No | Temporal graph memory |
Full API Reference
MemStack Client
import { MemStack } from "@memstack/core";
const ms = new MemStack({
llm: LLMProvider,
embedding?: EmbeddingProvider,
storage?: StorageProvider,
defaults?: {
summarizationThreshold?: number,
embedOnStore?: boolean,
pruneStrategy?: PruneStrategy,
},
hooks?: {
onMemoryStored?: (memory: Memory) => void;
onMemoryPruned?: (ids: string[]) => void;
onSummaryCreated?: (summary: Memory, deletedCount: number) => void;
},
});
Memory Subsystem
All methods accessible via ms.memory.*:
ms.memory.store(input: MemoryStoreInput): Promise<Memory>
ms.memory.storeBatch(inputs: MemoryStoreInput[]): Promise<Memory[]>
ms.memory.retrieve(query: MemoryRetrieveQuery): Promise<Memory[]>
ms.memory.get(id: string): Promise<Memory | null>
ms.memory.compileContext(options: ContextOptions): Promise<CompiledContext>
ms.memory.summarize(options: SummarizeOptions): Promise<{ summary: Memory; deletedCount: number }>
ms.memory.prune(strategy: PruneStrategy): Promise<{ pruned: string[]; count: number }>
ms.memory.dryRunPrune(strategy: PruneStrategy): Promise<{ wouldPrune: string[]; count: number }>
ms.memory.count(filter?: MemoryCountFilter): Promise<number>
ms.memory.delete(id: string): Promise<void>
ms.memory.deleteMany(ids: string[]): Promise<number>
ms.memory.touch(id: string): Promise<void>
Export / Import
Snapshot and restore full state for persistence, backups, or migration:
const snapshot = await ms.export();
fs.writeFileSync("state.json", JSON.stringify(snapshot, null, 2));
const data = JSON.parse(fs.readFileSync("state.json", "utf-8"));
await ms2.import(data);
Health & Close
const status = await ms.health();
await ms.close();
Configuration
const ms = new MemStack({
llm: new OpenAILLMAdapter({ apiKey: "..." }),
defaults: {
summarizationThreshold: 50,
embedOnStore: false,
pruneStrategy: {
type: "byAge",
maxAge: 90 * 86400000,
},
},
hooks: {
onMemoryStored: (m) => logger.debug("memory:stored", { id: m.id, actor: m.actorId }),
onMemoryPruned: (ids) => logger.info("memory:pruned", { count: ids.length }),
onSummaryCreated: (summary, n) => logger.info("memory:summarized", { count: n }),
},
});
Advanced Usage
Custom Storage
Implement StorageProvider for any database. The interface is 9 methods. See the reference section above for the full contract.
Custom LLM / Embedding
Implement LLMProvider or EmbeddingProvider for any service:
import type { LLMProvider } from "@memstack/core";
class TogetherAIAdapter implements LLMProvider {
async complete(req: { system: string; user: string; model?: string }) {
const res = await fetch("https://api.together.xyz/v1/chat/completions", {
headers: { Authorization: `Bearer ${this.apiKey}`, "Content-Type": "application/json" },
body: JSON.stringify({ model: req.model, messages: [{ role: "system", content: req.system }, { role: "user", content: req.user }] }),
});
const data = await res.json() as any;
return { text: data.choices[0].message.content, tokens: { prompt: data.usage.prompt_tokens, completion: data.usage.completion_tokens, total: data.usage.total_tokens } };
}
}
Event Hooks
Monitor memory operations without modifying code:
const ms = new MemStack({
llm,
hooks: {
onMemoryStored: (m) => metrics.increment("memory.stored"),
onSummaryCreated: (_, n) => metrics.gauge("memory.summarized_count", n),
onMemoryPruned: (ids) => metrics.increment("memory.pruned", ids.length),
},
});
Development
Setup & Tests
git clone https://github.com/isiomaC/memstack.git
cd memstack
pnpm install
pnpm test
pnpm test:watch
pnpm build
pnpm check
Debugging
Use hooks for observability — MemStack has no built-in logging:
const ms = new MemStack({
llm,
hooks: {
onMemoryStored: (m) => console.debug("[memstack] stored:", m.id, m.content.slice(0, 80)),
onMemoryPruned: (ids) => console.debug("[memstack] pruned:", ids.length),
},
});
Common issues:
CONFIG_ERROR: LLM provider is required | No LLM adapter | Pass any LLMProvider to config |
| Empty retrieval results | Wrong actorId or no memories stored | Check await ms.memory.count({ actorId }) |
| Semantic search not working | No embedding adapter or embedOnStore: false | Add embedding adapter or use strategy: "recent" |
| High memory usage in production | Using InMemoryStorageAdapter | Implement StorageProvider for Postgres/Redis/etc |
| Poor summarization quality | Default prompt doesn't match your domain | Use summarizationPrompt in defaults config |
Inspecting state at runtime:
const total = await ms.memory.count();
const perActor = await ms.memory.count({ actorId: "user-42" });
const snapshot = await ms.export();
const actorMemories = snapshot.memories.filter(m => m.actorId === "user-42");
console.log(`User-42: ${actorMemories.length} memories`);
actorMemories.forEach(m => console.log(` [${m.memoryType}] ${m.content.slice(0, 60)} (imp: ${m.importance})`));
Publishing to npm
pnpm build && pnpm check && pnpm test
npm login
npm publish --access public
The @memstack scope requires --access public on first publish.
Contributing
Most needed contributions:
- Docker Compose for integration testing
- LLM adapters: Google Gemini (native), Amazon Bedrock, Vertex AI
- Embedding adapters: local inference (transformers.js, ONNX)
- Benchmarks: retrieval quality, latency, cost comparisons
- Python port:
pip install memstack
Open an issue or PR at github.com/isiomaC/memstack.
License
MIT © MemStack