
Company News
Jerod Santo Joins Socket as Head of Media
Allow myself to introduce... myself.
Audit, test and measure the harness your AI agent runs on — grade your CLAUDE.md / AGENTS.md, skills, subagents and hooks, run them against a scripted model, and measure whether they actually fire.
You have a rule your agent follows half the time — and no way to know which one.
Verify your CLAUDE.md or AGENTS.md, skills, and hooks are real — and prove they actually work.
▶ Grade any public repo live at vigiles.sh — no install, runs in your browser.
You review every PR. Nothing reviews your CLAUDE.md.
Your harness — the CLAUDE.md or AGENTS.md rules, skills, subagents, and hooks steering your agent — is the one part nobody checks. Nobody verified it's real. Nobody tested it works. That's not a system. That's vibes.
And vibes break silently mid-task: a subagent wired to a tool that doesn't exist, two skills your agent can't tell apart, one helper quietly able to read your secrets and send them out.
vigiles1 checks your harness is real, not just well-formed — Claude Code and Codex alike. One command, no key, no config, safe on any repo:
npx vigiles audit
It's free and open-source, runs entirely on your machine, and never bills per token. (eval is the only step that calls a model — on your own Claude subscription.) Here's what it caught on plugins people actually ship. ↓
Like Google's Lighthouse, but for your agent harness. One command grades it A–F across six categories, leads with a plain-English verdict — "two one-line fixes away from a B" — and ranks every fix by the points it buys back:
0: if nothing asked, it says not measured.)And it closes the loop from prose to enforcement: your rules → enforced maps each rule you wrote to the lint rule that actually enforces it — already on, one config line away, or silently turned off (below).
These are real scans of public plugins — run npx vigiles audit <any-repo> for your own. The examples below use Claude Code subagents; the same checks run on Codex AGENTS.md, skills, and hooks. ↓
Every one of these is valid markdown — parses fine, does the wrong thing. That's the gap a style linter can't see:
"always use ===" while eqeqeq is set to "off" in your ESLint config. Your CLAUDE.md says enforce it; your config quietly turns it off..env anywhere — no exploit, just the tools it was given, hidden behind a healthy-looking grade.▶ See these live at vigiles.sh — grade any repo · everything it catches →. Point audit at a whole marketplace and it ranks every plugin the same way.
audit shows you where your setup is still vibes. Turning that into verified is four commands over one engine — and almost none of it needs a model or a key.
| Command | Answers | Needs a model? | When to run |
|---|---|---|---|
audit | Everything, graded A–F | No — read-only2 | Anytime; it's the report |
lint | Do the structural checks pass? | No | CI gate, every push |
test | Does the harness behave? | No — a scripted stand-in | Every commit |
eval | Does a skill actually help? | Yes — your subscription | On demand |
One engine, two doors. audit is the local report; lint is the CI gate that fails the build on the same deterministic checks — broken refs, bad tool contracts, dead hooks, skill collisions (Proofs 1–2). test and eval go further: past does it exist to does it work. (init / compile / eject manage the optional typed-spec layer for the structural rules no linter can express — a graduation step you rarely run by hand. If you use a spec, run compile in CI too: it is what re-derives that spec's refs, while lint verifies the compiled file is intact.) How the verbs relate →
Every path, script, symbol, and rule verified against reality — plus tool contracts, skill collisions, and dead hooks (the catches above). You don't write the checks, and you don't port anything into a spec: lint reads the CLAUDE.md or AGENTS.md you already hand-edit and checks it as-is. Rules that map to a real linter rule run in your own config (ESLint, Ruff), not a shadow layer.
How →
A hook that blocks nothing, a skill that hijacks unrelated prompts, context that never reaches the model — each passes a naive "did it run?" check. That gap is false confidence: a guard that looks like it works and silently doesn't. vigiles tests the real thing — hooks block, skills fire, subagents finish what they promised, a stray git push is caught before it happens. It drives a scripted stand-in for the model, not a live call, so it needs no key and runs on every commit.
How testing works → · Got a safety hook? Prove it blocks rm -rf and force-push →
Nothing you scan leaves your machine. lint, audit and the deterministic test tiers make no network call at all — no telemetry, no analytics, no HTTP client in the package. Evals drive your own claude CLI on your own subscription, so no third party is introduced. What is and isn't transmitted →
"Caveman Mode cuts 65% of your tokens." Says who? vigiles A/Bs the claim on real coding tasks and hands you three numbers: the token bill, whether it hit its target, and whether your code still works.
caveman vs baseline · sonnet · 7 tasks × 5 trials · $0 on your subscription
output tokens 6% lower on average (the claim was 65% — and it GREW on 2 of 7 tasks)
bill $0.5396 → $0.5334 (flat — output is only ~20% of the cost)
correctness 1.0 → 1.0 (nothing broke)
Point it at any harness change that claims a number — does a compression skill pay for itself, is a subagent worth its cost, which model is cheapest here. promptfoo and DeepEval bill per token, every run; vigiles runs on your own Claude Pro/Max subscription, so you measure on every change, not once. A committed lock file (like package-lock) keeps CI honest without re-calling the model. (Claude Code today; Codex landing.)
Measure a skill →
1. See what's broken — read-only, no setup (or try it in your browser first at vigiles.sh):
npx vigiles audit
2. Set it up when you like what you see. Paste into Claude Code or Codex:
Set up vigiles in this repo: run `npx vigiles init` and accept the defaults. If I
already have a CLAUDE.md or AGENTS.md, audit it and show me which references are
stale and which of my rules aren't enforced. Then write + run one harness test for
a hook or skill of mine. Don't run a real-model eval without asking me first.
Or run it yourself:
npx vigiles init # sets up the typed spec for structural rules (non-destructive — eject reverses), adds CI,
# installs vigiles's skills + hooks as a Claude Code plugin. The plugin
# CONTENT goes to ~/.claude/, never your repo; init also commits a two-key
# reference to it in .claude/settings.json so teammates get prompted to
# install it rather than silently missing it. Codex: skills install globally.
Already have a harness, or a non-JS repo? npx vigiles init --ci-only sets up just the CI integrity gate — nothing installed, zero conflict. When to use gate vs full →
Interactive in a terminal, non-interactive for agents/CI (or --yes). Works with Claude Code and Codex — vigiles verifies CLAUDE.md and AGENTS.md the same way. Codex setup →
Adoption is smooth: one command, then your agent does the rest. init installs the skills and hooks, so a plain-English ask does the work — no specs to hand-write, no hooks to wire:
test-harness)strengthen)edit-spec)The hooks keep it honest in-loop — nudging the agent to tag a linter-rule mention so vigiles can verify it, or to re-run a test whose result just went stale — so there are no chores to remember.
init sets up--lint / --test.audit and lint read them as-is — nothing is moved or rewritten. For the structural rules that want a typed spec, init sets one up non-destructively (eject undoes it).vigiles to devDependencies; installs the Claude Code plugin (skills + hooks) via the marketplace — globally, never vendored..claude/settings.json (extraKnownMarketplaces + enabledPlugins, merged into whatever is already there). This is a reference, not content — nothing is vendored, and it does not install the plugin for a teammate: an external-source plugin declared project-level doesn't load until each person installs it. What it buys is that Claude Code prompts them with the install command, instead of a fresh clone silently having the npm package and none of its skills.zernie/vigiles@v1 workflow (needs only read + PR-comment permissions) that posts a sticky PR comment + a valid output.Targets Claude Code and Codex out of the box, or your own harness. Prefer to write tests yourself? JS or TS (*.harness.{mjs,ts}) — run with npx vigiles test.
npm audit. One command, a report, an optional CI gate. There's a library API for automation, but you never touch it to get value.audit and lint read your local repo — no upload, no account, no server. (eval is the only step that calls a model, on your own Claude subscription; the browser demo only reads a public repo you name via GitHub's API.)strict (why?).npx vigiles lint verifies your CLAUDE.md or AGENTS.md with no install (Ruff/Clippy/golangci-lint/detekt/… too — 11 linters across Python, Rust, Go, Kotlin, Java, Ruby, CSS).Not for you if you want a model/capability benchmark or runtime guardrails in the request path — vigiles is build-/CI-time.
vigiles.sh is the live demo — grade any repo in your browser. The docs index is the full map, grouped by what you're doing:
A name starting with
experimental_is not covered by semver. It may change shape or disappear in a patch release. Everything so named has a stable alternative, given in that feature's own docs page — and the prefix is the only signal you need to look for, since it is on every call site rather than on an import line you scrolled past. See Stability.
Project — Stability · Contributing · Related tools · ships as an Agent Plugins 1.0.0 plugin (how to do the same)
vigiles — the watchmen of ancient Rome, who guarded the city (and fought its fires) by night. Quis custodiet ipsos custodes? — "who watches the watchmen?" (Juvenal, Satire VI). ↩
audit reads only by default. Two deeper checks — live MCP connections and skill-firing — are opt-in and ask before they run. ↩
FAQs
Audit, test and measure the harness your AI agent runs on — grade your CLAUDE.md / AGENTS.md, skills, subagents and hooks, run them against a scripted model, and measure whether they actually fire.
The npm package vigiles receives a total of 1,262 weekly downloads. As such, vigiles popularity was classified as popular.
We found that vigiles demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Company News
Allow myself to introduce... myself.

Research
/Security News
A Twitch browser extension on Chrome and Firefox forwards users’ live OAuth session tokens through proxies controlled by a Russian bot service.

Security News
Anthropic found biased reasoning and recklessness drove Claude Mythos 5 to publish malware on PyPI and compromise a security vendor.