vigiles
You have a rule your agent follows half the time — and no way to know which one.
Verify your CLAUDE.md or AGENTS.md, skills, and hooks are real — and prove they actually work.
▶ Grade any public repo live at vigiles.sh — no install, runs in your browser.
You review every PR. Nothing reviews your CLAUDE.md.
Your harness — the CLAUDE.md or AGENTS.md rules, skills, subagents, and hooks steering your agent — is the one part nobody checks. Nobody verified it's real. Nobody tested it works. That's not a system. That's vibes.
And vibes break silently mid-task: a subagent wired to a tool that doesn't exist, two skills your agent can't tell apart, one helper quietly able to read your secrets and send them out.
vigiles1 checks your harness is real, not just well-formed — Claude Code and Codex alike. One command, no key, no config, safe on any repo:
npx vigiles audit
It's free and open-source, runs entirely on your machine, and never bills per token. (eval is the only step that calls a model — on your own Claude subscription.) Here's what it caught on plugins people actually ship. ↓
What it caught
Like Google's Lighthouse, but for your agent harness. One command grades it A–F across six categories, leads with a plain-English verdict — "two one-line fixes away from a B" — and ranks every fix by the points it buys back:
- Truthfulness — do the references resolve?
- Triggering — do skills fire, without colliding?
- Structure — are tool contracts and configs valid?
- Safety — any way for the agent to leak your data?
- Tested — does the harness ship deterministic tests?
- Evaluated — has anything measured whether your skills actually fire? (Distinct from
0: if nothing asked, it says not measured.)
And it closes the loop from prose to enforcement: your rules → enforced maps each rule you wrote to the lint rule that actually enforces it — already on, one config line away, or silently turned off (below).
These are real scans of public plugins — run npx vigiles audit <any-repo> for your own. The examples below use Claude Code subagents; the same checks run on Codex AGENTS.md, skills, and hooks. ↓
Every one of these is valid markdown — parses fine, does the wrong thing. That's the gap a style linter can't see:
- A tool your agent thinks it has and doesn't — a subagent lists a tool that isn't available to it. The harness drops it silently; the agent loses a capability it believes it has.
- Two skills your agent can't tell apart — near-identical descriptions, so the agent (which picks by reading them) fires the wrong one. One popular plugin ships 45 such pairs.
- A rule you wrote that nothing enforces —
"always use ===" while eqeqeq is set to "off" in your ESLint config. Your CLAUDE.md says enforce it; your config quietly turns it off.
- A subagent that can read your secrets and send them out — one holding all three "lethal-trifecta" legs (reads data · takes untrusted web input · can send data out). A poisoned page makes it POST your
.env anywhere — no exploit, just the tools it was given, hidden behind a healthy-looking grade.
▶ See these live at vigiles.sh — grade any repo · everything it catches →. Point audit at a whole marketplace and it ranks every plugin the same way.
How it works — vibes → verified
audit shows you where your setup is still vibes. Turning that into verified is four commands over one engine — and almost none of it needs a model or a key.
audit | Everything, graded A–F | No — read-only2 | Anytime; it's the report |
lint | Do the structural checks pass? | No | CI gate, every push |
test | Does the harness behave? | No — a scripted stand-in | Every commit |
eval | Does a skill actually help? | Yes — your subscription | On demand |
One engine, two doors. audit is the local report; lint is the CI gate that fails the build on the same deterministic checks — broken refs, bad tool contracts, dead hooks, skill collisions (Proofs 1–2). test and eval go further: past does it exist to does it work. (init / compile / eject manage the optional typed-spec layer for the structural rules no linter can express — a graduation step you rarely run by hand. If you use a spec, run compile in CI too: it is what re-derives that spec's refs, while lint verifies the compiled file is intact.) How the verbs relate →
🔎 Lint — your instructions stop lying
Every path, script, symbol, and rule verified against reality — plus tool contracts, skill collisions, and dead hooks (the catches above). You don't write the checks, and you don't port anything into a spec: lint reads the CLAUDE.md or AGENTS.md you already hand-edit and checks it as-is. Rules that map to a real linter rule run in your own config (ESLint, Ruff), not a shadow layer.
How →
🧪 Test — does the harness actually do its job?
A hook that blocks nothing, a skill that hijacks unrelated prompts, context that never reaches the model — each passes a naive "did it run?" check. That gap is false confidence: a guard that looks like it works and silently doesn't. vigiles tests the real thing — hooks block, skills fire, subagents finish what they promised, a stray git push is caught before it happens. It drives a scripted stand-in for the model, not a live call, so it needs no key and runs on every commit.
How testing works → · Got a safety hook? Prove it blocks rm -rf and force-push →
Nothing you scan leaves your machine. lint, audit and the deterministic test tiers make no network call at all — no telemetry, no analytics, no HTTP client in the package. Evals drive your own claude CLI on your own subscription, so no third party is introduced. What is and isn't transmitted →
📊 Eval — the only way to put a real number on cost
"Caveman Mode cuts 65% of your tokens." Says who? vigiles A/Bs the claim on real coding tasks and hands you three numbers: the token bill, whether it hit its target, and whether your code still works.
caveman vs verbose · haiku · $0 on your subscription
output tokens 762 → 842 (+11% — the "saving" reversed)
correctness 1.0 → 1.0 (the fact survived)
Point it at any harness change that claims a number — does a compression skill pay for itself, is a subagent worth its cost, which model is cheapest here. promptfoo and DeepEval bill per token, every run; vigiles runs on your own Claude Pro/Max subscription, so you measure on every change, not once. A committed lock file (like package-lock) keeps CI honest without re-calling the model. (Claude Code today; Codex landing.)
Measure a skill →
Quick start
1. See what's broken — read-only, no setup (or try it in your browser first at vigiles.sh):
npx vigiles audit
2. Set it up when you like what you see. Paste into Claude Code or Codex:
Set up vigiles in this repo: run `npx vigiles init` and accept the defaults. If I
already have a CLAUDE.md or AGENTS.md, audit it and show me which references are
stale and which of my rules aren't enforced. Then write + run one harness test for
a hook or skill of mine. Don't run a real-model eval without asking me first.
Or run it yourself:
npx vigiles init
Already have a harness, or a non-JS repo? npx vigiles init --ci-only sets up just the CI integrity gate — nothing installed, zero conflict. When to use gate vs full →
Interactive in a terminal, non-interactive for agents/CI (or --yes). Works with Claude Code and Codex — vigiles verifies CLAUDE.md and AGENTS.md the same way. Codex setup →
Adoption is smooth: one command, then your agent does the rest. init installs the skills and hooks, so a plain-English ask does the work — no specs to hand-write, no hooks to wire:
- "test my skills" → scaffolds and runs a trigger/behaviour test, then commits its result so CI can check it (
test-harness)
- "harden my rules" → upgrades prose guidance into enforced linter rules (
strengthen)
- "add a rule to my CLAUDE.md or AGENTS.md" → edits the source and recompiles (
edit-spec)
The hooks keep it honest in-loop — nudging the agent to tag a linter-rule mention so vigiles can verify it, or to re-run a test whose result just went stale — so there are no chores to remember.
What init sets up
- Both lint and test by default; scope with
--lint / --test.
- Already have a CLAUDE.md / AGENTS.md, skills, or subagents?
audit and lint read them as-is — nothing is moved or rewritten. For the structural rules that want a typed spec, init sets one up non-destructively (eject undoes it).
- Adds
vigiles to devDependencies; installs the Claude Code plugin (skills + hooks) via the marketplace — globally, never vendored.
- Declares that plugin in your repo's
.claude/settings.json (extraKnownMarketplaces + enabledPlugins, merged into whatever is already there). This is a reference, not content — nothing is vendored, and it does not install the plugin for a teammate: an external-source plugin declared project-level doesn't load until each person installs it. What it buys is that Claude Code prompts them with the install command, instead of a fresh clone silently having the npm package and none of its skills.
- Wires CI as a
zernie/vigiles@v1 workflow (needs only read + PR-comment permissions) that posts a sticky PR comment + a valid output.
Targets Claude Code and Codex out of the box, or your own harness. Prefer to write tests yourself? JS or TS (*.harness.{mjs,ts}) — run with npx vigiles test.
FAQ
- Is this a framework I have to build around? No. It's a tool you run — like ESLint, Lighthouse, or
npm audit. One command, a report, an optional CI gate. There's a library API for automation, but you never touch it to get value.
- Isn't this just a markdown linter? No — it checks whether your instruction file is true (every path/script/symbol/rule exists and is enabled), then tests and measures your harness. A style linter can't do any of that.
- Does anything leave my machine? No.
audit and lint read your local repo — no upload, no account, no server. (eval is the only step that calls a model, on your own Claude subscription; the browser demo only reads a public repo you name via GitHub's API.)
- What do I have to change to adopt it? Almost nothing. Plain markdown works with zero new files, rules run in your existing linter, and the agent edits the specs for you. The typed spec is opt-in, only for structural checks a linter can't express — like TS's
strict (why?).
- Non-JS repo?
npx vigiles lint verifies your CLAUDE.md or AGENTS.md with no install (Ruff/Clippy/golangci-lint/detekt/… too — 11 linters across Python, Rust, Go, Kotlin, Java, Ruby, CSS).
Full FAQ →
Not for you if you want a model/capability benchmark or runtime guardrails in the request path — vigiles is build-/CI-time.
Docs
vigiles.sh is the live demo — grade any repo in your browser. The docs index is the full map, grouped by what you're doing:
A name starting with experimental_ is not covered by semver. It may change
shape or disappear in a patch release. Everything so named has a stable
alternative, given in that feature's own docs page — and the prefix is the only
signal you need to look for, since it is on every call site rather than on an
import line you scrolled past. See Stability.
Project — Stability · Related tools · ships as an Agent Plugins 1.0.0 plugin (how to do the same)
License
MIT