
Security News
Re-Enabled GitHub Actions Expose Thousands of Repositories to Mini Shai-Hulud
Two compromised GitHub Actions were re-enabled with malicious tags intact, exposing thousands of downstream repositories to Mini Shai-Hulud.
A temporal knowledge engine for markdown vaults — indexes, links, remembers, forgets, and dreams. Local-first, agent-ready (MCP).
A temporal knowledge engine for markdown vaults.
It indexes, links, remembers, forgets, and dreams — locally, over files you own.
Quickstart · Different how? · CLI · Agent memory · How it works · Research
Most knowledge tools are write-only. You capture diligently, the vault grows, and six months later you can't find the thing you know you wrote — because retrieval is keyword search over prose, nothing ever resurfaces on its own, and nothing notices when what you wrote last year stopped being true.
Loreweave is the layer that fixes that. Point it at a folder of markdown (Obsidian or plain) and it builds a knowledge graph, a bitemporal fact store, and a memory model over your notes — then hands them to you through a CLI and to your AI agents through MCP.
Your files stay exactly as they are. The vault is the source of truth; the index is a cache you can delete at any time.
npx loreweave init && npx loreweave index
cd ~/my-vault
npx loreweave init # creates .lore/
npx loreweave index # incremental; ~1.2 ms per note (see Scale below)
npx loreweave search "why did we drop the queue design"
npx loreweave ask "what's the status of project atlas"
npx loreweave dream # what's duplicated, contradicted, stale, unlinked
Zero configuration required and no network calls: out of the box it runs on BM25 + knowledge-graph spreading activation. Add embeddings when you want them:
// .lore/config.json
{ "embedding": { "provider": "ollama", "model": "nomic-embed-text" } }
Everything degrades gracefully — no embedding provider means lexical + graph retrieval, still fully functional.
Works with non-English vaults: Chinese, Japanese and Korean text is segmented per character so it is searchable at all, and other scripts index as written.
Loreweave ships its own benchmark — npm run eval — over three purpose-built
vaults, scored against a BM25 baseline and a graph-only baseline.
The second corpus exists to catch overfitting: it is deliberately unlike the
first in every dimension the config could have been tuned to (markdown links
instead of [[wiki links]], deep nesting, filenames unrelated to titles, real
engineering note shapes). The third measures time instead of topicality:
facts change across dated notes, every date lives only in frontmatter where
BM25 cannot see it, and each windowed question is paired with a
shifted-window twin whose correct answer is a different note — the
perturbation test that exposes systems faking temporal competence through
lexical overlap.
| corpus | system | finds the answer | in top 5 | MRR | answer shown |
|---|---|---|---|---|---|
| kestrel (40 q) | hybrid | 100% | 75% | 0.545 | 55% |
| BM25 | 75% | 65% | 0.532 | 55% | |
| northwind (24 q) | hybrid | 96% | 92% | 0.690 | 83% |
| BM25 | 63% | 58% | 0.521 | 54% | |
| meridian (18 q) | hybrid | 100% | 94% | 0.952 | 94% |
| BM25 | 94% | 83% | 0.437 | 83% |
Multi-hop is where the graph earns its keep: BM25 finds 0% on both prose
corpora — it cannot reach a note that shares no words with your query, at
any depth — while hybrid finds 90% and 100%. Time is where the temporal
machinery earns its keep: on meridian's windowed questions hybrid ranks the
right note first 100% of the time vs BM25's 0% — and on the paired
perturbation test (same question, shifted window, different correct answer,
both directions must rank first) hybrid scores 100% vs BM25's 0%. The
consistency number is computed and regression-gated by npm run eval, not
hand-derived.
"answer shown" is the strictest measure: not just the right note, but a returned passage that literally contains the answer. Results are one per note, showing whichever of that note's sections best covers your query — ranking decides which notes matter, coverage decides which part of them you see.
The same shipped config wins on all three corpora, and by more on the ones it
was never tuned against. If you prefer pure lexical behaviour, set
retrieval.weights.expansion: 0.
Run it yourself: npm run eval. npm run eval:gate fails the build on any
regression across all three corpora, and CI enforces it on every push.
What lexical + graph retrieval cannot do, stated precisely: the kestrel multi-hop questions use words like "hardware" and "outpost" that appear in zero notes — the vault says "instrument" and "station". No statistic derived from the vault (co-occurrence, PPMI, LSA) can bridge that, because there are no occurrences to derive one from. Those answers are still found by following links; ranking them first needs semantics from outside the vault.
And embeddings do not currently supply it — measured. npm run eval -- --embed reruns everything with local dense vectors (Ollama +
nomic-embed-text) on top. The result on these corpora is worse, not better:
| corpus | r@5 | MRR | answer shown | |
|---|---|---|---|---|
| kestrel | lexical + graph | 0.750 | 0.545 | 0.550 |
| + embeddings | 0.725 | 0.539 | 0.525 | |
| northwind | lexical + graph | 0.917 | 0.690 | 0.833 |
| + embeddings | 0.833 | 0.604 | 0.792 | |
| meridian | either | 0.944 | 0.951 | 0.944 |
Sweeping the dense fusion weight (1.0 → 0.5 → 0.25 → 0) does not recover it,
and the loss is split between the dense ranking list and the similarity edges
it adds to the graph. So the honest position: the numbers above are what the
shipped, model-free configuration does, and we cannot claim the optional layer
improves anything. The likely reason is these corpora themselves — invented,
distinctive vocabulary is exactly where lexical matching is strongest and
paraphrase is rarest. On prose full of synonyms it may well pay off; measure it
on your own vault with --embed rather than trusting either of us.
The corpora above are ours — we wrote the notes and the questions, which makes them good for regression and worthless as proof. These are third-party, with relevance labels nobody here chose. All runs are the model-free configuration (no LLM, no embeddings, no network), and all are retrieval metrics: loreweave finds the evidence, it does not write the answer, so these are not comparable to end-to-end QA accuracy quoted by systems that put a language model after retrieval.
BEIR / SciFact — 5 183 abstracts, 300 queries, nDCG@10:
| system | nDCG@10 | Recall@10 |
|---|---|---|
| loreweave, lexical channel | 0.682 | 0.808 |
| loreweave, full pipeline | 0.676 | 0.811 |
| BM25 (BEIR paper, Anserini) | 0.665 | — |
LongMemEval_S (ICLR 2025) — 500 questions, ~50-session history each, session-level recall:
| category | n | R@1 | R@5 | R@10 |
|---|---|---|---|---|
| knowledge-update | 78 | 0.481 | 0.942 | 0.955 |
| temporal-reasoning | 133 | 0.414 | 0.871 | 0.910 |
| multi-session | 133 | 0.371 | 0.835 | 0.925 |
| single-session (all) | 156 | 0.833 | 0.955 | 0.974 |
| overall | 500 | 0.552 | 0.899 | 0.943 |
LoCoMo — 10 long conversations, 1 982 evidence-labelled questions, turn-level recall: R@1 0.340, R@5 0.539, R@10 0.612, R@20 0.655. Its strongest category is temporal (R@1 0.451), its weakest multi-hop (R@1 0.087) — a question needing four turns can score at most 0.25 at R@1.
Two things worth saying plainly. None of these corpora have links, tags, or frontmatter — the structure this engine exists to exploit — so it is being measured with one hand tied; that is the honest cost of using benchmarks we didn't design. And measuring them found a real defect: the graph channel was trusted even on corpora with no graph to walk, which cost 0.024 nDCG@10 on SciFact and 6 points of R@5 on LoCoMo. Fixed in 0.32.0. It is not uniformly better — LongMemEval's R@10 moved 0.955 → 0.943 — and the trade is documented rather than hidden.
Reproduce every number: docs/benchmarks.md.
Measured, like the quality numbers — npm run scale reproduces this on your
own machine (synthetic vaults, 3 blocks per note, dense interlinking):
| notes | blocks | entities | edges | full index | incremental | search p50 | p95 | dream | heap |
|---|---|---|---|---|---|---|---|---|---|
| 1 000 | 3 000 | 2 766 | 17 k | 1.2 s | 32 ms | 3 ms | 4 ms | 0.2 s | 63 MB |
| 5 000 | 15 000 | 13 766 | 85 k | 6.1 s | 165 ms | 9 ms | 12 ms | 1.0 s | 145 MB |
| 20 000 | 60 000 | 55 016 | 339 k | 22.5 s | 656 ms | 38 ms | 53 ms | 4.8 s | 333 MB |
Full index scales at 0.93× per note from 5 k to 20 k — linear or better,
no superlinear step hiding in the middle. "Incremental" is one changed note,
which is what lore watch actually does all day. Everything here is one
process, one SQLite file, no daemon.
1. Knowledge that has a timeline. Facts are bitemporal: when they were true in the
world (valid_from/valid_until) and when the system learned them (recorded_at).
Contradictions supersede rather than overwrite, so history stays queryable.
$ lore assert "Ledger Format" status draft --valid-from 2026-01-01
$ lore assert "Ledger Format" status final --valid-from 2026-08-01
✓ Ledger Format :: status :: final
superseded: "draft" (now valid until 2026-08-01)
journal: lore/journal/2026-08-01.md
$ lore facts --subject "Ledger Format"
Ledger Format :: status :: final (2026-08-01 → now)
asserted · lore/journal/2026-08-01.md
$ lore facts --subject "Ledger Format" --as-of 2026-03-01
Ledger Format :: status :: draft (2026-01-01 → 2026-08-01) [superseded]
asserted · lore/journal/2026-08-01.md
Both axes are queryable, which is the part that makes it bitemporal rather than
merely historical. --as-of asks what was true then; --as-known-at asks what
was believed then, excluding anything recorded later however far back it was
backdated. They disagree exactly when you learn something after the fact — which
is when you most need to reconstruct what a past decision was actually based on:
$ lore facts --subject Vendor --as-of 2024-06-01
Vendor :: reliability :: poor — outage postmortem (2024-01-01 → now)
$ lore facts --subject Vendor --as-known-at 2024-06-01
Vendor :: reliability :: good (2024-01-01 → 2024-01-01) [superseded]
Which fact wins is decided deterministically (newest valid-time, provenance as tiebreak) — never by asking a language model which one looks fresher.
And the whole history of anything is one command — every value change from the fact store merged chronologically with the dated prose that mentions it:
$ lore timeline Project Atlas
2024-01-15 status: planning (until 2024-09-01)
2024-02-10 • [[Project Atlas]] kicked off with a three-person crew. [kickoff.md]
2024-09-01 status: planning → active
2025-06-20 • The [[Project Atlas]] midpoint review went long but well. [review.md]
"What was X before it changed" is the query temporal-graph products market as their flagship — built there by running an LLM over every ingested document. Here the supersede chain has been maintained all along, so it is a read-side join: no LLM, no network, same answer every time.
2. Retrieval that follows connections, not just words. Queries fuse BM25, dense similarity (when configured), and Personalized PageRank over the vault's own graph — wiki-links, shared entities, tags, co-occurrence. Two-hop neighbors surface even when they share no vocabulary with your query, and every result tells you why:
• data/glacier-dataset.md#@0 (0.0327) ⟨via amara osei⟩
The Glacier Dataset holds meltwater sensor readings from 2019-2024.
3. Memory with dynamics. Every passage carries FSRS-style stability and retrievability — a power-law forgetting curve. Passages that actually get used (not merely retrieved) decay slower; important-but-fading knowledge gets surfaced for review instead of silently rotting. Nothing is ever deleted.
4. It dreams. lore dream is an idle-time consolidation pass that reviews the vault
and reports duplicate passages, contradicted facts, stale knowledge, missing links
between notes that clearly belong together, and orphans. With --apply it writes a
digest and a review queue — append-only, under lore/. It never rewrites your prose:
LLM-driven whole-file rewriting is a documented failure mode (context collapse), so the
architecture forbids it.
5. Questions retrieval can't answer. Counting, grouping, and date-range queries run as deterministic SQL over the fact store, not as vibes over embeddings:
$ lore count --predicate trip_to --since 2025-01-01 --until 2025-12-31
2 Japan
1 Kenya
6. Facts come from your notes, not from a form. The fact store used to be
empty on any real vault — nobody hand-writes - [fact] X :: y :: z. It now
mines the conventions vaults already use:
status: shipped # frontmatter
- owner:: Priya # Dataview inline field
- [location] Hyderabad # Basic Memory observation
Only unambiguous field syntax is accepted automatically. Prose formatting like
- **Owner:** Priya is precise on entity notes and noisy on report notes, so
it is opt-in (facts.extract: "all") — or an agent can review candidates via
lore_propose_facts and assert the real ones. Judgement stays out of the index.
7. Time means when it happened, not when you saved the file. --since and
--until filter on content time — taken from frontmatter dates, dated
filenames (2025-03-14-standup.md), or dates in the text — falling back to
file mtime only when a note carries no date of its own:
$ lore search ledger --since 2025-01-01 --until 2025-12-31
• 2025-03-14-standup.md › Standup [all terms]
Discussed the ledger migration.
Both files were written seconds ago, so an mtime filter could not tell them
apart. lore watch keeps the index current so you never have to remember to
reindex.
8. Built for agents. An MCP server exposes 15 typed tools so Claude Code, Cursor, or
any MCP client can use your vault as durable memory — with a session context pack,
fact assertion, point-in-time queries, and a reinforcement signal. Session
continuity is a query, not a paraphrase: lore resume returns exactly what
changed since the agent last connected, computed from record time —
$ lore resume
since 2026-08-11 15:55
~ lore/journal/2026-08-11.md
+ Project Atlas :: status :: shipped (since 2026-08-11)
± Project Atlas :: status: active → shipped
$ lore resume
since 2026-08-11 15:56
nothing changed
— where the popular alternatives run an LLM over the previous session and inject the summary: a paraphrase, unreproducible, wrong exactly when it matters.
| Command | What it does |
|---|---|
lore init | create .lore/ with a default config |
lore index [--full] [--no-nlp] [--rebuild-similar] | incremental sync of vault → index |
lore search <q> [-k] [--since] [--until] [--tag] [--folder] [--json] | hybrid retrieval with provenance |
lore ask <q> | extractive answer: current facts + top passages (no LLM needed) |
lore facts [--subject] [--predicate] [--as-of] [--as-known-at] [--history] | query the fact store |
lore timeline <entity> [--since] [--until] | chronological history: fact changes merged with dated mentions |
lore resume [--since] | what changed since the last resume: notes, facts, supersessions |
lore review [--threshold] [--limit] | important-but-fading knowledge to revisit or archive |
lore assert <s> <p> <o…> [--valid-from] | record a fact (journalled, supersedes) |
lore invalidate <s> <p> | close the current fact in a slot |
lore count [--predicate] [--group-by] [--since] | aggregate over fact history |
lore capture <text…> | append a timestamped line to lore/inbox.md |
lore dream [--apply] | consolidation pass + optional digest/review queue |
lore watch | reindex automatically as the vault changes |
lore mark-used <note> [anchor] | reinforce a passage that proved useful |
lore graph export --format json|graphml|dot | export the graph |
lore doctor | health check: broken links, integrity, coverage |
lore stats | vault statistics and top entities |
lore serve --mcp | start the MCP server on stdio |
// Claude Code: .mcp.json (or claude_desktop_config.json)
{
"mcpServers": {
"loreweave": {
"command": "npx",
"args": ["-y", "loreweave", "--vault", "/path/to/vault", "serve", "--mcp"]
}
}
}
Tools: lore_search, lore_context_pack, lore_read_note, lore_assert_fact,
lore_invalidate_fact, lore_query_facts, lore_timeline, lore_resume, lore_review, lore_aggregate_facts, lore_capture,
lore_mark_used, lore_propose_facts, lore_dream_report, lore_index.
Facts asserted through MCP are written back to lore/journal/YYYY-MM-DD.md as readable
markdown lines, so an agent's memory is something you can open, read, edit, and
git diff:
- [fact] Ledger Format :: status :: final {valid_from=2026-08-01, confidence=0.9, source=stated}
Delete .lore/ and reindex — every fact and edge is reconstructed from those files.
vault/*.md ──parse──▶ notes · blocks · wiki-links · tags · entities
│ (incremental: mtime + content hash)
▼
SQLite .lore/index.db ── disposable cache, rebuildable
│
┌───────────────────┼────────────────────┐
▼ ▼ ▼
graph (CSR) retrieval facts
blocks ∪ entities BM25 + dense + PPR bitemporal, supersession,
2-iteration PPR → weighted RRF deterministic freshness,
α = 0.5 → FSRS boosts aggregates
└─────────┬─────────┴──────────┬─────────┘
▼ ▼
dream (idle-time) CLI · MCP
Design rules the code enforces:
lore/.Every significant choice traces to 2024-2026 literature; the full 87-finding survey lives
in docs/research/ and the reasoning in
docs/superpowers/specs/.
| Choice | Source |
|---|---|
| Dense-sparse fusion + PPR with dense reset probabilities | HippoRAG 2 (ICML 2025), 2502.14802 |
| Shallow 2-iteration PPR, heterogeneous nodes | NodeRAG (2025), 2504.11544 |
| Relation-free graph — no LLM triple extraction | LinearRAG (ICLR 2026), 2510.10114; AtomicRAG (2026) |
| No index-time community summarization | LazyGraphRAG (Microsoft, 2024) — same quality at 0.1% index cost |
| Route/fuse instead of graph-everything | GraphRAG-Bench (ICLR 2026), 2506.05690 |
| Bitemporal facts, invalidate-never-delete | Zep/Graphiti (2025), 2501.13956 |
Typed version links (updates/extends/derives) | Supermemory, SOTA on LongMemEval |
| Deterministic freshness, not LLM-judged | "Don't Ask the LLM to Track Freshness" (2026) |
| Power-law forgetting, use-gated reinforcement | FSRS; RMM (ACL 2025), 2503.08026 |
| Consolidation as idle-time work | Sleep-time compute (Letta, 2025), 2504.13171 |
| Never let an LLM rewrite whole memory files | ACE (2025), 2510.04618 |
| Computable facts for aggregation | User as Code (2026), 2606.16707 |
| Fine-grained indexing + fact-augmented keys | LongMemEval (ICLR 2025), 2410.10813 |
import { openContext, indexVault, search, assertFact, queryFacts, dream } from 'loreweave';
const ctx = openContext('/path/to/vault');
await indexVault(ctx.store, ctx.root);
const hits = await search(ctx, 'streaming compaction', { k: 5 });
assertFact(ctx, { subject: 'Atlas', predicate: 'status', object: 'shipped', validFrom: '2026-08-01' });
const asOfMarch = queryFacts(ctx.store, { subject: 'Atlas', asOf: '2026-03-01' });
const report = dream(ctx);
ctx.close();
npm install
npm test # 421 tests
npm run eval # retrieval benchmark vs BM25 baseline
npm run typecheck
npm run build
Requires Node ≥ 20. Single native dependency (better-sqlite3). Tested in CI on
Linux, macOS and Windows across Node 20 and 22.
MIT © Ambuj Upadhyay
FAQs
A temporal knowledge engine for markdown vaults — indexes, links, remembers, forgets, and dreams. Local-first, agent-ready (MCP).
The npm package loreweave receives a total of 45 weekly downloads. As such, loreweave popularity was classified as not popular.
We found that loreweave demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Security News
Two compromised GitHub Actions were re-enabled with malicious tags intact, exposing thousands of downstream repositories to Mini Shai-Hulud.

Research
/Security News
A malicious Firefox extension fetches its payload after installation to evade detection, steal Google session cookies, and automate account takeover.

Research
/Security News
The compromise affects MemTensor's MemOS, an open source memory framework for large language models (LLMs) and AI agents. Both npm package @memtensor/memos-cloud-openclaw-plugin and the PyPI package MemoryOS are compromised. They drop cross-platform Go binaries that exfiltrate developer secrets.