
Security News
/Company News
Securing the Financial Frontier: How Capital One Uses Socket for Open Source Security
Capital One is partnering with Socket to proactively secure its open source supply chain.
@agentguard-run/burn
Advanced tools
Local session usage and runaway-agent circuit breaker. Explain tokens, cache rewrites and API list cost, track pace, warn on heavy turns, and gate agent fan-out with signed receipts. Developer command ranking requires explicit --rank.
See where a coding session's tokens went and what the next heavy turn could cost in time and limits. Burn explains recorded usage, warns about large cache rewrites and keeps its existing runaway-session circuit breaker. Runtime observation and enforcement stay on this machine. The developer review command below has an explicit external-ranking option.
npm i -g @agentguard-run/burn
agentguard-burn init claude
agentguard-burn init codex
agentguard-burn statusline
agentguard-burn why
The two init commands print hook configuration for review. Merge the relevant snippet into your host configuration. They do not write it. For Codex, review and trust the hook through /hooks. Upgrading an existing hook requires the new snippet: it matches every tool so pace can update on ordinary tool calls, while admission still gates spawns.
agentguard-burn --version (or -v) prints the installed version.
For Claude Code, configure its statusLine command as agentguard-burn statusline. It receives the host JSON on stdin. Codex 0.154.0 has no external status command slot. Use agentguard-burn statusline SESSION_ID in a companion terminal and the native Codex limit indicators. See host fields and setup.
Burn 0.3.3 can verify and replay a signed DisciplineBench bundle offline. This is a post-run display, outside benchmark scoring. It does not run an agent or change a benchmark result.
agentguard-burn import-bench ./run-bundle --out ./burn-import
agentguard-burn replay ./burn-import/replay.json --pricing ./burn-import/bundle/pricing.json
agentguard-burn replay ./burn-import/replay.json --pricing ./burn-import/bundle/pricing.json --json
agentguard-burn render ./burn-import/recording.jsonl --mp4 ./run.mp4 --gif ./run.gif
The importer uses the bundled, unchanged DisciplineBench verifier to check the manifest signature, every declared artifact hash, the Ed25519 ledger chain, and the recomputed run.json. Only manifest-listed artifacts are copied into a new directory. The source bundle is unchanged. Existing output directories are refused. The default output is the bundle path with .burn appended.
Embedded keys prove integrity, not the publisher's external identity. Obtain a publisher fingerprint through a trusted channel and supply --publisher-key-id ID to import or replay when that identity matters. Publisher and measurement fingerprints remain separate. No AgentGuard license, provider credential, model call, or network request is required.
Replay requires --pricing on every invocation. Its bytes must match the signed round's pricing artifact exactly. Current Burn prices, local overrides and prices from another round are never substituted. Cache durations and long-context rates use the harness pricing rules. The display rounds the final list-price exposure to cents; JSON keeps the unrounded reducer value. List price is a comparability convention, never a bill.
The replay file keeps every canonical event, the complete spawn tree, rule IDs, enforced:false, all native and normalized token categories, and the usage reconciliation. Missing measurements remain null and render as unknown. A STOP in this recording is a would-STOP classification, not an execution block. Burn does not invent a spawn ceiling or an active window absent from the bundle.
Each event produces a frame at its original timestamp. Frames contain cumulative observations through that event, not a synthetic 15-minute enforcement window. The final frame agrees with the verified reducer. replay ... --record frames.jsonl also writes those event frames. render uses the existing local PNG, MP4 and GIF pipeline and records source SHA-256 provenance. Frames are display derivatives, not newly signed measurements; replay re-verifies the retained bundle and checks the conversion before producing new frames.
Imports are limited to 128 MB of declared artifacts and 10,000 events, matching the local renderer's frame limit. No task code or modules from the bundle are executed. The exact vendored verifier sources, Apache-2.0 license and source hashes are in vendor/disciplinebench/.
The bundled fixtures/bench-run is an authored fixture, not a scored host run. Its reducer and Burn both report $0.0155, displayed as $0.02. It declares two child agents, maximum depth two, and a would-STOP observation for GP001 and GP006. See the rendered frame, MP4, GIF, and artifact hashes. The original harness's burn-postrun.json marker remains unchanged as signed evidence, including its older unsupported-import message.
guard-candidates reads local ledgers and writes a review file. It never executes a recorded command or changes a rule, policy, hook, or ledger.
agentguard-burn guard-candidates --out ./commands.review.json
agentguard-burn guard-candidates --ledger ./local-ledger.ndjson --out ./explicit.review.json
The default scan covers all named decision, receipt and ledger NDJSON/JSONL files under AGENTGUARD_HOME (default ~/.agentguard), plus PLUGIN_DATA and CLAUDE_PLUGIN_DATA when provided. Numbered/date suffix rotations are included. Nested symlinks are not followed. Use repeatable --ledger to choose other local ledger files. The command reads each file up to its opening size, across every session, including a complete last line without a newline. It does not read transcripts or shell history.
Commands are deduplicated by exact text, with occurrence counts. Whitespace, quoting, multiline programs and long strings stay intact. Recognized shell tool records can carry command, cmd or script in tool input, including signed decision.plugin envelopes and JSON function arguments. Non-shell tools, responses and incidental command fields are excluded.
Current Burn and plugin ledgers omit command text. They store metadata and input hashes. Those shell records are counted as missing text; hashes cannot reconstruct commands. This command does not enable new content logging. A metadata-only ledger produces an empty review, not invented candidates. Malformed rows are reported by file and line; ranking refuses an incomplete scan.
To rank, first install the optional jsort CLI, distributed as jev-sort under MIT. The adapter targets 0.1.3 (reviewed source). Burn does not install it automatically.
uv tool install 'jev-sort==0.1.3'
# Set TYPESAFE_API_KEY or OPENROUTER_API_KEY in your environment first.
agentguard-burn guard-candidates --ledger ./local-ledger.ndjson --out ./ranked-input.review.json --rank
Only --rank with a nonempty environment key can start jsort or a provider request. TypeSafe takes precedence when both keys are present. Key files and gateway/model overrides do not enable ranking. With no key, or fewer than two unique commands, the raw review is kept and ranking is skipped.
The exact criterion is how likely is this command to destroy data, leak a secret, escalate privilege, or change infrastructure. Before jsort starts, the command prints the provider, endpoint, criterion and complete command texts as escaped JSON. Each exact outgoing JSON request body is then printed and flushed before it is forwarded. Authentication keys are never printed or stored in review files. Recorded command text is sent verbatim, including any secrets inside it, so inspect the local review before choosing --rank.
jsort receives only command strings through stdin. A local forwarding gate permits only comparisons between those strings, using that criterion and the selected provider/model. It rejects unexpected request data and redirects. The provider key stays in Burn; jsort gets a temporary local token. No cache is written. The adapter uses jsort's $1 budget, a 15-second comparison timeout and a five-minute process timeout. Incomplete, failed or malformed rankings leave the raw review intact and return a nonzero exit status.
Successful ranking creates ranked-input.ranked.review.json beside the raw file, with each command's score, standard error and comparison count, highest first. Scores are relative Bradley-Terry logit values, not calibrated risk probabilities. All outputs have owner-only permissions. Existing outputs, non-review output names and --record are refused. These commands have no automatic connection to Guard Pack rules and are never installed or called by hooks.
Burn 0.3.0 reports held memory and disk cost exposure on this machine:
agentguard-burn ps
agentguard-burn ps --json
agentguard-burn reap
ps lists Claude Code and Codex sessions with their Node children, leftover plugin daemons, automation-owned Chrome browsers, and candidate workspaces. Process rows are ranked by resident memory, then workspace rows by disk size. Normal Chrome profiles are excluded. Browser rows show their process count, total resident memory and available window count. Unknown values are explicit. RSS can count shared memory more than once; it is not a bill.
Session idle time uses the controlling terminal's access time when readable, otherwise transcript file-write metadata. The table labels the method. A transcript fallback can warn but cannot prove that a terminal has been inactive for 24 hours. Workspace rows show size, last modification, dirty status and open-handle evidence. No transcript content or process arguments are printed or uploaded. Swap is reported only: swap is only released by the operating system on reboot. Burn does not touch it.
reap prints the same numbered audit, then requires an interactive terminal and numbers typed at the prompt. There is no --yes flag and piped confirmation is refused. It sends only SIGTERM to the selected eligible processes, waits, and reports whether they exited. It never sends SIGKILL and never deletes files or directories.
Reap refuses:
It rechecks the selected processes before signaling. Permission errors and processes that ignore SIGTERM are reported. They are not force-closed.
Add these fields under thresholds in the existing ~/.agentguard/burn-policy.json (or AGENTGUARD_HOME/burn-policy.json):
{
"thresholds": {
"idle_session_warn_hours": 24,
"idle_browser_warn_minutes": 60,
"orphan_workspace_warn_days": 2
}
}
Merge the fields into your existing policy. Missing or invalid values use the shipped defaults. A local teamPolicyFile can point to a JSON file containing thresholds; relative paths resolve from the Burn home. Team threshold values take precedence over local threshold values, including these three audit settings. Neither policy loading nor the audit rewrites your policy.
The Claude Code and Codex hook snippets now include SessionStart. Review the new agentguard-burn init claude or agentguard-burn init codex output to update existing installations. At session start, a fresh five-minute local audit cache is checked. A cache miss gets a bounded scan with a 150 ms budget; timeout or unavailable evidence skips the advisory. A warning is one line, never a table and never a deny:
N idle agent sessions holding X GB, run agentguard-burn ps
The line is WARN-only in both shadow and enforce modes. If only browser or workspace rows cross their thresholds, the session count can be zero; run ps for those details. The full ps command refreshes the cache. All collection and caching stays local.
macOS uses local process, file-handle and swap metadata. Linux uses /proc and available local tools, and says what could not be read. Windows prints not supported on this platform yet and exits successfully. See the docs page draft for the audit boundary and limitations.
agentguard-burn rewrites
agentguard-burn rewrites SESSION_ID
agentguard-burn rewrites all
agentguard-burn rewrites SESSION_ID json
A full-prefix rewrite writes more than 150,000 tokens and more than half the actual input context. Every rewrite gets one row: when it happened, the session, the host, the model, the cause, the idle gap before it, the cache-write tokens, the list-price equivalent of that write and a one-line explanation. The by-cause summary follows, and the last line is the rewrite total as a share of the same session list total that why prints. Without a session, rewrites reads the newest top-level session plus its stored child transcripts, exactly as why does.
Causes are named in a fixed order, so a rewrite with several signals gets the first one that applies: model_switch (the model differs from the previous usage turn), version_upgrade (Claude Code's recorded version differs), compaction, first_spawned_turn, idle_ttl_expired (the idle gap exceeds the lifetime recorded for the write), image_pruning (an earlier message was re-recorded with fewer image or PDF blocks), then a host-reported prefix_change. What remains is prefix_change as a residual, or unknown_ttl when the lifetime is absent or mixed. Codex transcripts carry the model on turn_context and mark compaction with compacted; a Codex rewrite with neither signal is unknown, with the reason stated, never prefix change by default.
Three causes in Claude Code's cache-invalidation list are not written to the transcript, so Burn never names them: effort changes, fast mode and tool definition changes (MCP connect or disconnect, plugin toggles, denying a whole tool). Claude Code also does not persist image pruning; the image_pruning rule only fires when a message is re-recorded without its images. The explanation column says which of these could not be observed instead of guessing.
The synthetic fixture in tests/fixtures/rewrite-fixtures.ts renders as follows (rows shortened):
AgentGuard rewrites: 10; 1 followed an idle gap over 60 minutes.
When Session Host Model Cause Idle Write tokens List ...
2026-09-24T09:20:00Z synthetic-claude-fixture claude claude-sonnet-5 idle_ttl_expired 20m 180,000 $0.45 ...
2026-09-24T09:22:00Z synthetic-claude-fixture claude claude-opus-5-5 model_switch 2m 180,000 $1.44 ...
2026-09-24T09:24:00Z synthetic-claude-fixture claude claude-opus-5-5 version_upgrade 2m 180,000 $1.44 ...
2026-09-24T09:26:00Z synthetic-claude-fixture claude claude-opus-5-5 compaction 2m 180,000 $0.90 ...
2026-09-24T12:26:00Z synthetic-claude-fixture claude claude-opus-5-5 idle_ttl_expired 3h 180,000 $1.44 ...
...
Rewrites: $9.47, 95.6% of the combined session list total $9.90 (API list equivalent, not a bill).
The JSON form is stable: {sessions: [{sessionId, host, records: [{at, model, cause, explanation, idleMs, cacheWriteTokens, usd}]}], totals: {count, cacheWriteTokens, usd, shareOfSessionUsd}}. Rewrite tokens are always a subset of the measured cache-creation tokens; nothing is estimated from text.
agentguard-burn ttl
agentguard-burn ttl SESSION_ID
agentguard-burn ttl all
agentguard-burn ttl SESSION_ID json
ttl answers one question from recorded turns: had the cache lifetime been one hour instead of five minutes, what would the same history have been billed at API list rates? A rewrite counts as avoided when it followed an expired five-minute lifetime with an idle gap of at most one hour; under a one-hour lifetime those tokens would have been billed as cache reads. Against that, every other write recorded with a five-minute lifetime would have been billed at the one-hour write rate. The report keeps the main conversation and subagent transcripts in separate buckets because Claude Code sets their lifetimes separately. Unknown lifetimes stay unknown and are excluded, with the excluded count shown.
The same fixture renders as:
AgentGuard ttl: synthetic-claude-fixture, agent-synthetic-child
List-price counterfactual on recorded turns: the result is the net at list price, not a bill. Subscription users: read the token columns; every dollar column is the API list equivalent.
Bucket Rewrites avoided Avoided write tokens Billed at 5m Would have been reads Repriced 5m write tokens Extra at 1h rate Net at list
main 1 of 8 180,000 $0.45 $0.04 751,000 $2.07 -$1.66 (5m)
subagent 1 of 2 170,000 $0.42 $0.03 160,900 $0.24 +$0.15 (1h)
Rewrites avoided: full-prefix writes that followed an expired five-minute lifetime with an idle gap of at most one hour. Under a one-hour lifetime they would have been billed as cache reads, not writes.
Extra at 1h rate: every other write recorded with a five-minute lifetime would have been billed at the one-hour write rate instead. Net = billed at 5m minus reads minus extra.
main: 1 rewrite with an unknown lifetime excluded from the counterfactual
Main conversation: net favours 5m. Leave promptCacheTtl unset; the longer lifetime would have been billed more on these turns.
Subagent transcripts: net favours 1h. Set "subagentPromptCacheTtl": "1h" in settings.json.
On a Claude subscription within plan usage the main conversation already has a one-hour lifetime; API key and usage-credit sessions start at five minutes. Subagents, forks, compaction and titles stay at five minutes unless subagentPromptCacheTtl is set.
Multipliers relative to base input: 5m write 1.25x, 1h write 2x, cache read 0.1x; a model row's verified rates take precedence.
Every dollar column is an API list equivalent; on a subscription, read the token columns. The multipliers come from Anthropic's prompt caching pricing (five-minute writes 1.25x base input, one-hour writes 2x, cache reads 0.1x) and only fill rates a model row leaves empty; the verified rates in the pricing table, and any local override, take precedence. The verdict lines name the exact setting: "promptCacheTtl": "1h" for the main conversation (or CLAUDE_CODE_PROMPT_CACHE_TTL=1h) and "subagentPromptCacheTtl": "1h" for subagents, forks, compaction and titles. On a Claude subscription within plan usage the main conversation already has a one-hour lifetime; API key and usage-credit sessions start at five minutes. When the net favours five minutes, the report says so plainly.
agentguard-burn why
agentguard-burn why SESSION_ID
agentguard-burn pace SESSION_ID
agentguard-burn pricing
A session argument can also be a transcript path. Without one, Burn uses the session environment variable when available, then the most recently modified local transcript. Claude child transcripts are included in why. Token shares are shares of recorded tokens, not shares of dollars. Repeated usage records for one provider response count once. Copied Codex history is reconciled before counting new usage.
The table covers instruction stack + system, history re-sent, repeated file reads, subagent fan-out, tool output, conversation, full-prefix rewrites, output and unattributed usage. The fixed-prefix baseline is the first assistant response's measured context. It includes the initial user message, so the label is an operational baseline rather than a direct measurement of instruction files. A rewrite classified as prefix change resets that baseline and is counted in the footer; the new baseline can include history already present at that point.
The report ends with at most one tip: the one action that matters most for this session, from its own numbers. A cause stands out at these shares: sub-agents at 20% or more of session tokens, full-prefix rewrites at 15% or more of the session's API list equivalent (skipped when a model is unpriced), and the conversation Claude re-reads, or the setup it reads before every message, at 30% or more of session tokens. Among the causes that stand out, the largest at API list price wins (the largest in tokens when a model is unpriced), and the tip names one action from Anthropic's own guidance (costs, prompt caching, sub-agents):
38 sub-agents ran on Opus 5.5. and a smaller model for searching and reading logs (CLAUDE_CODE_SUBAGENT_MODEL=haiku, or model: haiku in the sub-agent's definition). Otherwise the sub-agent share, and AgentGuard's free plugin, which asks you before the next sub-agent past your limit starts./compact before a break. Reloads from switching models: how many, and switching at a break right after /compact. Any other reload cause: the dollars and agentguard-burn rewrites./compact at a natural break or /clear between unrelated tasks./mcp).The setup and the re-read conversation split each call's cached reads at the size of the session's first call, so a rewrite that resets the table's baseline never turns re-read conversation into setup. Tips never claim a saving and print no dollar figure except the rewrite dollars the table prints. When nothing stands out there is no tip. agentguard-burn quiet on hides it, quiet off shows it again and quiet says which. why json never contains it.
After a /compact, an automatic compaction, a model switch, or a /clear that the transcript records, why adds one measured line before the tip, for example Measured change: after /compact on 2026-09-27 at 14:02 UTC, Claude read a median of 85K tokens per message, down from 660K in the 20 messages before. It compares the median tokens per message of the main conversation in the N messages after the latest such event with the N before it: N is 20, fewer when the session or the next or previous event leaves fewer, and with fewer than 5 on either side that event says nothing. It states a measurement next to an event, never a saving or a cause, and prints whether or not tips are on.
why prints in Latin American Spanish or Brazilian Portuguese when your machine is set to that language, and --lang es, --lang pt or --lang en picks one explicitly. The order is --lang, then AGENTGUARD_LANG, then LC_ALL, LC_MESSAGES and LANG (the first one set decides, as POSIX says, so LC_ALL=C means English), then the locale Node resolves. Every other language is English. Figures keep their en-US formatting in every language (1,234,567, $1335.37, 12.3%), so a translated report matches screenshots digit for digit. why, compare and calendar are translated. Their json output is the same in every language, and every other command stays English.
Later cache reads up to the baseline go to instruction stack + system. Reads above it go to history re-sent. Fresh input plus cache creation is the measured arriving increment. Burn subtracts the previous response's output before assigning that increment to intervening tool results and user messages. Result byte sizes only split that measured total; bytes are never converted into tokens. A result for a previously read Read path goes to re-read files. Mixed user and tool intervals are marked shared and counted in the footer.
Rewrite tokens and child transcript usage keep their own buckets. Prior output deducted from an input increment remains in the unattributed residual so every recorded input and output token is counted exactly once. Missing event evidence also remains there. Integer allocation preserves every token, and the Method footer explains each row.
Every tool hook updates a local pace file under the Burn home. Pace uses ten wall-clock minutes: cached means cache reads; uncached means fresh input, writes and output. The next-hour projection assumes that pace continues. A time-to-limit estimate requires rising, fresh host percentage samples from the same quota pool and reset window. No token-to-limit conversion or cache weighting is assumed. Without sufficient host data, the line says why the limit is unknown.
The new warnings are advisory in both modes. Cache rewrites warn once when their last-hour API list equivalent exceeds $5. A context crossing 500,000 tokens warns about processing time and compaction, estimates a cache rewrite, and suggests a fresh session with a handoff note. These warnings do not deny a tool call.
Optional settings in your existing burn-policy.json:
{
"insights": {
"rewriteWarnDollarsPerHour": 5,
"heavyTurnTokens": 500000
}
}
Merge these fields into the existing policy. All existing thresholds and modes remain supported. Prices are exact-model API list equivalents, not a subscription bill. Unknown models show tokens only. Every rate has a source and verification date. A local override file and full pricing table are documented in Usage and pricing.
agentguard-burn compare
agentguard-burn compare SESSION_ID
agentguard-burn compare SESSION_A SESSION_B
agentguard-burn compare SESSION_A SESSION_B json
Run the same task once on each model, then put the two sessions side by side. With no session, compare takes the two most recent top-level sessions in the project of the newest one; with one, that session and the newest other top-level session in its project; with two, those two in the order given. Sessions resolve as they do for why (id, id suffix or transcript path). Sub-agent transcripts and other files Claude Code keeps beside a session are never picked. A Codex rollout's project is the working directory it recorded, and a spawned sub-agent rollout is skipped. Unless you name both, A is the session that started first. Fewer than two sessions exits 1 with one line saying how to pass two ids.
Both sessions are named first: the short id (the last part of the session id, which compare and why accept back), the time of the first recorded response in UTC, the model with the most output tokens (every model with at least 5% of output tokens, with its share, when there are several) and the project path. Each figure is the one why prints for that session, child transcripts included. Change is B against A, only where both values are numbers and A is not zero as printed; a dollar range compares midpoints, and an unpriced model prints unpriced with no change. The fixture in tests/fixtures/compare-fixtures.ts renders as:
AgentGuard compare: B against A
A aaaaaaaaaaaa started 2026-09-25 10:01 UTC
Output by model: claude-opus-5-5 92.3%, claude-haiku-4-5-20251001 7.7%
Project: /synthetic/project
B bbbbbbbbbbbb started 2026-09-25 11:01 UTC
Model: claude-sonnet-5
Project: /synthetic/project
Metric A B Change
Assistant responses 4 3 -25.0%
Input tokens 599,600 543,000 -9.4%
read from cache 200,000 360,000 +80.0%
cache writes 396,000 180,500 -54.4%
fresh input 3,600 2,500 -30.6%
Cache share 33.4% 66.3% +98.8%
Output tokens 6,500 4,000 -38.5%
Sub-agents launched 1 0 -100.0%
Full-prefix rewrites 2 1 -50.0%
List USD $2.13 $0.57 -73.3%
Totals include sub-agents.
Dollar figures are API list-price equivalents, not a bill.
AgentGuard cannot tell whether these sessions did the same task. Compare runs of the same task.
compare measures and does not rank models. --lang es|pt and your locale choose the language exactly as for why. compare json is never translated: both sessions with ids, paths, project, models and every metric, and each change as an unrounded fraction of A.
agentguard-burn calendar
agentguard-burn calendar --weeks 12
agentguard-burn calendar json
Every Claude Code and Codex session on this machine on one heatmap, sub-agents included: a column per week, a row per weekday, darker for more tokens. Each response's tokens and API price come from the parser and prices why uses, counted on the local calendar day the response was recorded. A response recorded in two files (a sub-agent's copy of its parent's last reply, a project that moved, a forked session) counts once, merged the way why merges a session. A forked Codex session whose copied history has no verifiable boundary is left out and named in a note, as why reports it unavailable.
The summary is three short lines, however wide the terminal: the total tokens and the dollars at API prices (not a bill; from $1,000 in whole dollars, such as $28,252 to $28,254 at API prices, not a bill, and one number when both ends round to the same dollar); the busiest day and the share of tokens sub-agents used (every token of their responses); and, only when there were any, how many sub-agent launches AgentGuard held: refused by an enforced STOP, or held for your yes, as the decisions ledger records them. On a narrow terminal a line's two items take a line each. A day is amber when sub-agents used more than half of its tokens, and carries a red mark when AgentGuard held a launch that day. The five shades run from Less to More: a day's level is the quarter of the shown days with tokens that it reaches, so the busiest day shown is always the darkest. Tokens on a model without a verified rate are counted but left out of the dollars, and a note names the model.
With no flag the calendar starts at the oldest recorded day, capped to the weeks that fit the terminal (its width on a terminal, 80 columns when piped, as when Claude Code runs ! npx agentguard-burn calendar); --weeks N shows exactly N weeks, 1 to 53, ending with this week.
Days are big whenever the span fits: two characters a day and one space between weeks, which reads as roughly square. That takes the weekday labels plus three columns a week, so 25 weeks fit 80 columns, 39 fit 120 and all 53 fit 162. When the span does not fit (a long history at 80 columns, or a --weeks span wider than the terminal), each day is one character and one space, as 0.3.20 drew every calendar; the span never shrinks to make the days bigger. A colour terminal draws a big day as ▇▇ and a compact day as ■. The seven-eighths block leaves a thin gap above each day, where full blocks would join a week's days into one bar. Piped, with --no-color or with NO_COLOR set, a big day is a shade pair (·· ░░ ▒▒ ▓▓ ██) and a compact day one shade character (· ░ ▒ ▓ █). The two kinds of marked day are the letters s and H (the red diamond in colour): a big day keeps its shade in the left half and shows the mark in the right (▓s, █H), while a compact day shows the mark alone. The fixture in tests/fixtures/calendar-fixtures.ts prints four weeks, in big days:
AgentGuard calendar: tokens per day, Aug 30 to Sep 26, 2026
696,040 tokens · $7.29 at API prices, not a bill
busiest day: Sep 24 (511,040 tokens) · sub-agents used 7.3% of tokens
AgentGuard held 2 launches
Sep
Su ·· ·· ··
Mo ·· ▓s ··
Tu ░░ ·· ·· ··
We ·· ▒▒ █H ··
Th ·· ·· ·· ██
Fr ·· ·· ·· ··
Sa ·· ·· ·· ▒▒
Less ·· ░░ ▒▒ ▓▓ ██ More
Days marked s: sub-agents used more than half of that day's tokens.
Days marked H: AgentGuard held a launch that day.
From Claude Code and Codex sessions on this machine, sub-agents included.
Claude Code deletes sessions older than 30 days by default.
Burn keeps each day's count whenever you run this, so the calendar grows.
Next: see where one session's tokens went: npx agentguard-burn why
The last line is the one next step, chosen the way why chooses its tip: when sub-agents took a fifth or more of the tokens it points at the free plugin (or, once AgentGuard has held launches on this machine, at giving related work to one sub-agent); otherwise it points at why, which names the change for one session.
Claude Code deletes transcripts after 30 days by default (cleanupPeriodDays), so each run keeps a daily count per transcript file in calendar/daily.json under the Burn home (AGENTGUARD_HOME, default ~/.agentguard): tokens, sub-agent tokens and dollars per day, with no paths and no content, owner-only and written atomically. A transcript still on disk is always counted from the transcript; a deleted transcript's days come from the daily count; no file-day is counted from both. Run calendar at least once within Claude Code's retention period to keep every day. calendar/cache.json keeps each response's counts, keyed by the transcript's path digest, size and modification time, so a later run re-reads only the transcripts that changed; a new Burn version or a pricing override recounts everything once.
calendar json is never translated: every day of the span with its figures and level, the totals, the busiest day and where the figures came from. --lang es|pt and your locale choose the language exactly as for why. Nothing in the calendar compares you with anyone else.
agentguard-burn card
agentguard-burn card --weeks 1 --out this-week.png
One 1200 by 675 PNG with the calendar's totals, made to post: the token total with the Claude Code and Codex split, the dollars at API prices (not a bill), the share sub-agents used, the busiest day, the sub-agent launches AgentGuard held (or the transcripts read, when it held none) and one bar per day in whole weeks. --weeks N takes 1 to 53 and defaults to 4. The image is saved as agentguard-burn-card.png in the current folder unless --out names another file. It is drawn on your machine from the same figures calendar prints; no project, prompt or file name appears on it, and nothing is uploaded. A span with no recorded usage is refused.
Structural limits use active time rather than the age of the session. The lifetime spawn count warns at 24 but cannot stop a session on its own. The fan-out ceiling permits 40 spawns within the last 120 active minutes and stops the next proposal. The separate spawn-rate rule warns at 8 and stops candidate 16 within 15 active minutes. Depth above 2 still stops.
Economic limits warn at 3.5B and stop at 5B measured tokens per session. Repeated Claude assistant content blocks share one provider response id. Their usage and copied child history count once, including later increases to a response's recorded usage. Old session counters rebuild once from the parent and its available child transcripts, retaining the signed chain head. The rebuild places events chronologically. A child usage record discovered late on a later hook enters the current active-minute bucket, while session and bucket totals remain conserved. The same numeric sustained thresholds now fire later because inflated usage has been removed. Earlier replay figures in the changelog used the old counts.
Existing policy files keep explicit settings. Missing fields are filled in
memory, with a one-time notice. In particular, an old explicit
spawnRate.enforce: false remains advisory until you change it. New policies
default to true; fanout.windowActiveMinutes defaults to 120. Run calibration
only when you want it to write a new shadow policy, or set AGENTGUARD_HOME
to a scratch directory to review fitted values first.
agentguard-burn blocks
agentguard-burn blocks SESSION_ID
agentguard-burn blocks all
Each WARN or STOP receipt gets one row with its time, session, detector reasons and spawn ordinal. The row includes the session's available child count, median and maximum tokens, and median and maximum API list-price dollars per child. Forks and fresh agents are reported separately. The report reuses the attribution parser and excludes copied provider responses. It reads recorded usage only and makes no network calls.
These are list-price equivalents, not a bill or a claim of money saved.
Economics describe the currently available child transcripts of that session,
not only the children completed before the receipt. Missing child history or
unknown model pricing stays unavailable. Unknown cache lifetimes produce a
price range. A fork is identified by its transcript's fork metadata. Codex child discovery
requires an explicit parent thread id; unknown inherited-history boundaries
stay unavailable. Use blocks all json for structured output. Session
attribution uses local receipt identifiers without copying transcript content
into the ledger.
agentguard-burn limits
agentguard-burn limits SESSION_ID
agentguard-burn limits all
agentguard-burn price-shadow
agentguard-burn price-shadow SESSION_ID
agentguard-burn price-shadow all
Both commands are read-only. They read the existing local host transcripts and
Burn ledgers, use no network or data plane, and do not load or change policy,
create receipts, or record frames. A session can also be a local transcript
path. Append json for structured arithmetic.
limits reports Claude Code plan-window rejections and allowed warnings, plus
HTTP 429 rate-limit errors. Each event shows its recorded time, window, reset,
model when present, and the closest earlier Burn WARN or STOP in that session.
The rollup counts rejected events, not allowed warnings. Reset durations use
the union of overlapping rejection-to-reset intervals, including across
sessions, so retries and shared windows do not multiply minutes. This is the
time until the recorded reset, not a measurement of time unable to work. An
absent reset stays unknown. Fan-out means a fanout, spawn_rate or depth finding
within the preceding 60 minutes. Proximity is correlation. Codex limits remain
unsupported until a verified transcript format is available.
price-shadow prices every explicitly recorded shadow decision using the
shipped API list snapshot, including OK and WARN rows. It includes only children
whose recorded start is strictly after that decision. Claude starts come from
the child's own sidechain user event or a timestamped fork marker; Codex starts
come from session metadata with explicit parent lineage. Existing children
that continue running after a decision do not enter its cohort. Copied parent
responses are excluded; provider response ids reconcile streaming updates and
copies across child transcripts. Medians and maxima use each child's full
available usage. Unknown starts, missing transcripts, unresolved spawn results
and unknown inherited-history boundaries are unpriced. Nothing is imputed.
Decision cohorts overlap. Totals use only the first would-block point per session and label the resulting list-price exposure an upper bound. A known priced subset is separate from an unpriced total. Fable 5.1 and Sonnet 5 reprice the same recorded token categories as sensitivity, never measurement. Unknown cache-write lifetimes produce a range. Dollars are list-price equivalents, not the bill. Neither command changes enforcement.
Cache-read ratio was about 98% in the original calibration, healthy and pathological alike. It is shown as an explanation and never used to decide.
npx @agentguard-run/burn replay
Runs the detectors over your existing history and shows what enforcement would have intercepted, when, and the observed tail after each stop. It is an upper bound, labelled as such. Nobody installs a blocker cold.
npm i -g @agentguard-run/burn
agentguard-burn init --write # merge the hook into ~/.claude/settings.json (backup taken)
agentguard-burn init codex --write # same for ~/.codex/hooks.json
agentguard-burn status # hook health, shadow observations, every live session
agentguard-burn enforce # after 7 days and 50 decisions
The hook installs in shadow mode: every decision is recorded, nothing is
blocked, until you have seen it be right. status shows what it would have
done so far, which sessions are live on every host, and whether each hook's
command still exists on disk (a hook whose script is gone fails open, and
status says so in capitals).
In an interactive Claude Code session (Manual, accept edits, auto or plan mode) an enforced STOP holds the launch behind Claude Code's own permission prompt. The prompt names the one limit that fired, for example:
AgentGuard: 15 sub-agents in the last 15 active minutes. Limit: 15.
Session so far: 15 sub-agents, 154.0M tokens.
The 12 sub-agents that finished in this session averaged 10.0M tokens each ($4.99 at API prices, not a bill). This one could differ.
Allow this one launch? If you say no, nothing starts.
The third line appears once at least three of the session's sub-agents have finished: what they averaged, measured from their own transcripts as the hook reads them anyway (never a second pass), and never a forecast. With fewer, the prompt is the other three lines. The box of a refused launch carries the same line.
You answer with one keypress. Claude never sees that text, and auto mode
cannot approve it silently (Claude Code 2.1.211 and later), so the answer is
yours. Where no prompt can appear (print mode, bypassPermissions, dontAsk,
Codex, Cursor) the launch is refused with the STOP box instead, and the box
ends with the override to run. In Claude Code, type it with ! in front so it
runs in your own shell:
! npx agentguard-burn resume --once --reason "these 60 agents are the plan" # the next STOP passes, once, on any host
! npx agentguard-burn resume --reason "load test" # every STOP passes for 15 minutes
! npx agentguard-burn resume --clear
Only the person can override. A ! command never passes through a hook, so
AgentGuard refuses the same resume, shadow or calibrate command when the
agent runs it, and refuses agent writes to override.json or
burn-policy.json. Set "stopStyle": "deny" in burn-policy.json to refuse
outright even where a prompt could show.
Every override is written to the decisions ledger with its reason. A
--once override is consumed atomically: two hooks racing for it cannot
both pass. Warnings are spoken once per change in the finding set, not once
per spawn; eighteen identical banners train you to stop reading the
nineteenth.
The detectors never learn which host produced an event. Claude Code, Cursor,
Codex, a local model runtime behind the proxy, and an orchestrator calling
the middleware all normalise into the same event stream, share one
machine-wide reservation lock, and sign the same receipt. The same failure,
through every door, stops at the same step: agentguard-burn conformance
replays a 42-spawn storm and a 250M-token-per-call grind through each adapter
and asserts fan-out WARN at 24, STOP at 41, sustained WARN at 3.5B, STOP at
5B.
What each host can actually see is stated, not implied:
| Host | Spawns | Depth | Usage | How |
|---|---|---|---|---|
| Claude Code | authoritative | authoritative | authoritative | PreToolUse hook + transcript (unchanged from 0.1) |
| Raw middleware | authoritative | authoritative | authoritative | beforeSpawn / beforeCall leases in your orchestrator |
| Ollama proxy | none | none | authoritative | prompt_eval_count + eval_count on the final chunk |
| vLLM / LM Studio / OpenAI-compatible proxy | none | none | authoritative when the server sends usage, else reported missing | non-streaming usage, or the final SSE usage event |
| Cursor (beta) | authoritative | estimated | none | native subagentStart deny; hosted-model usage is never exposed |
| Codex (beta) | authoritative | estimated | estimated | PreToolUse on every tool, with admission on spawn_agent; live deny, allow and override canary passed on codex-cli 0.151.0; transcript parsed best-effort |
An OK from a host that cannot see usage is an OK about spawns, and status
says usage:n/a next to it. Missing usage never becomes a guessed zero.
The full claim, "40 spawns, depth 2, 5B tokens, enforced identically", is true for a deployment that feeds both a topology source and a usage source into one session ID: raw middleware plus the proxy, for instance. A proxy alone sees tokens and no tree. A Cursor hook alone sees the tree and no tokens. The composite conformance check proves the combined case: candidate spawn 41 sees both planes in its findings.
Token dollars are close to meaningless when the GPU is yours. What runs away
is the machine: concurrency and occupied request time. The proxy tracks both
and warns at 4 concurrent calls by default. No universal STOP ships for
hardware we cannot see; set localCompute.stopConcurrent or
stopOccupiedMs in burn-policy.json for your server. Elapsed request time
includes queueing and transport, so it is called occupied time, never GPU
utilisation.
agentguard-burn proxy --upstream http://127.0.0.1:11434 --host ollama
# point the agent at http://127.0.0.1:18080 and send x-agentguard-session: <id>
Loopback only, both sides, by default. Every upstream chunk is written to the client before it is inspected; the observer is a side channel, never a data path. A STOP answers the next request with 429 and the alarm box. It never cuts a stream that is already flowing, and it never kills a running agent. Blocking is not killing.
import { createRawApiGuard } from '@agentguard-run/burn/middleware';
const burn = createRawApiGuard({ sessionId: 'nightly-refactor-17' });
const spawn = burn.beforeSpawn({ parentDepth: 0 });
spawn.throwIfBlocked();
spawn.started();
try { await worker() } finally { spawn.finished() }
const call = burn.beforeCall({ estimatedTokens: 120_000 });
call.throwIfBlocked();
try {
const res = await client.chat({ ..., headers: call.headers }); // proxy correlates by call id
call.complete({ tokens: res.usage.total_tokens });
} catch (e) { call.fail(); throw e }
Usage is committed by call ID and replaces what was reserved under it. When middleware estimated 120K and the proxy later saw 87K for the same call, the session moves by 87K, not 207K.
agentguard-burn init cursor # ~/.cursor/hooks.json snippet, failClosed on
agentguard-burn init codex # ~/.codex/hooks.json snippet
Both renderers are one page each and emit only their host's documented output
object. Codex fails the whole hook on Claude's common fields (continue,
stopReason, suppressOutput), and a failed hook is a fail-open hook, so the
Codex renderer never emits them and a test forbids them by name.
Codex passed a live canary on the installed codex-cli 0.151.0 on
2026-09-03: a spawn_agent call was denied with the alarm box as the reason,
a shell call passed untouched, and a resume --once override let the next
spawn through with the reason on the ledger. Two things the docs got wrong
and the wire settled: the tool arrives as spawn_agent (the docs say it
"matches as Agent"; the matcher covers both), and project-local hooks only
load when the project is trusted. The captured payloads are in
fixtures/codex-0.151.0-pretooluse.json and drive a test.
One step Codex makes you do by hand: after init codex --write, open
codex, run /hooks, and trust the AgentGuard hook. Codex requires this
once per hook source, and there is no CLI for it. Until it is done,
codex exec stalls the first time a spawn would be gated (verified: it
hangs with no output; the same with hooks defined in config.toml).
status repeats this under the codex line because it cannot see whether
trust was granted.
Cursor is verified against the documented schema (permission,
user_message, agent_message; ~/.cursor/hooks.json with failClosed),
not yet against an installed build. It stays beta until a live deny canary
passes.
Every spawn decision, and every model call that is not OK, is signed with a local Ed25519 key (Node built-ins, key generated on first use, 0600) and chained to the previous receipt for the session. A receipt carries the host, the coverage, the counts, the verdict, the policy digest and a hash of the session ID. It carries no prompt, completion, path, or tool input. It can leave the machine when a transcript never can.
Ten parallel Agent calls launch ten hook processes that all read the same
transcript and all see the same count. A naive cap is cosmetic during exactly
the burst it exists for. Spawns are admitted through an atomic, cross-process
reservation under a machine-wide lock; the test suite launches 60 real OS
reservation processes against a cap of 40 and asserts exactly 40 are admitted,
and separately exercises 60 Cursor hook processes. The opt-in stress test runs
240 actual Claude hook processes against cap 40:
npm run build
AGENTGUARD_STRESS=1 node --test dist/tests/stress.test.js
It checks exactly 40 admitted and 200 denied, verifies the receipt chain for coordinated decisions, and reports lock failures separately. Infrastructure failures deny in enforce mode before a receipt can be chained, as in the gateway.
0.2.0 fixed the lock itself. Under 240 concurrent hook processes the 0.1 lock could tear down a live sibling's lock (a waiter judged "owner is dead" about an instance that had already been released and replaced) and admit 41 to 46. Lock instances now carry a nonce; a reclaim only counts if it grabbed the instance it judged, and every write is fenced on the holder's own nonce still being on the path. In 0.2.3, bounded, staggered retries prevent lock waiters from starving the holder. The Sep 15 audit includes one successful 240-process run after that fix; it does not establish a 192-run guarantee. In 0.2.4, the wait is eight seconds and retirement is serialized before renaming an abandoned lock. The fresh 240-process run admitted exactly 40, denied 200, signed 240 receipts and had zero lock failures. Operational lock-wait failures do not count toward enforce eligibility. See the conservative recovery procedure.
Single-machine by design. Two laptops on one account do not share state, and that is stated rather than hidden.
account.warnConcurrentSessions warns when more than that many sessions have
been active on this machine in the last 30 minutes, matching status liveness.
Optional account.warnTokens and account.stopTokens sum the last
account.windowActiveMinutes of token buckets from each active session.
Both token limits default to null. Closed gateway sessions are excluded;
hook and gateway session files are both read. WARN never blocks. STOP records
a would-block decision in shadow and blocks new work only in enforce mode.
account.stopConcurrentSessions (default null) is a hard ceiling on live
sessions for the machine. Above it, every spawn proposal on the machine is
refused until a session finishes; other tool calls are untouched, and a plain
status evaluation only warns. It is meant for hosts that start sessions on
their own, such as Claude Code Projects running threads locally: the project
can run as many threads as it likes, but none of them may add copies while
the machine is over your number.
Interactive screens use a fixed 104 by 35 cell canvas, with mint for good, slate for neutral, amber for WARN and red for STOP. Smaller terminals drop side detail and paginate rows without wrapping. Use --no-color for the existing plain-text commands and hook cards. Hook JSON and the one-line statusline remain host protocols, with no canvas control sequences inserted.
Any command accepts an explicit local recording path:
agentguard-burn ps --record ./idle-run.jsonl
agentguard-burn status --no-color --record ./status-run.jsonl
agentguard-burn replay --record ./history-run.jsonl
agentguard-burn render ./idle-run.jsonl --mp4 ./idle-run.mp4 --gif ./idle-run.gif
Each JSONL frame contains a wall-clock timestamp and structured observations, never rendered terminal text. Audit recordings omit raw process arguments and the full process table. They can contain local paths, PIDs, identifiers and resource measurements, so review them before sharing. Recording is opt-in and stays on disk. A recording failure prints a notice without changing a hook decision or cleanup result. Appended recordings must have nondecreasing timestamps.
The render command draws recorded Canvas cells at 1920 by 1080 using installed Menlo on macOS or DejaVu Sans Mono on Linux. The title says RECORDED RUN and the footer says recorded run. It preserves frame timestamp intervals and holds the final frame for one second. ffmpeg encodes the MP4 and optional GIF; without it, numbered PNG files and a reproducible ffmpeg command remain. No fonts or assets are fetched. The native PNG dependency loads only for render, never during session-start checks.
A provenance JSON beside the requested MP4 records the source recording's SHA256, renderer version, render time, font and frame times. Existing outputs are refused. The source recording stays unchanged. See the render draft for the format and limits.
Runtime observation and enforcement do not persist or render prompts, responses, file contents or tool inputs. The explicit developer guard-candidates command copies shell text already present in a ledger into a private review; only its --rank option can send that text to the selected provider.
No telemetry. No provider-quota guesses. Account forecasts use only observed host percentages, never an invented allowance or cache weighting.
agentguard-burn live follows the session with the newest local ledger decision
on every refresh. --session ID pins a full ID, unique ID prefix or digest.
The header shows the raw session ID prefix when the decision ledger supplies it.
Spawn and token counters, the rate ceiling and the last eight decisions come
from that session's decision ledger. Pending spawn proposals are included,
exactly as the hook measured them. Elapsed starts at its first ledger decision.
A change of session is captured as sessionSwitch in --record run.jsonl.
Use --once, --record run.jsonl, or --replay run.jsonl. STOP reasons hold
for two seconds while observations continue, and a session switch releases
that hold immediately. The display never changes admission. Historical token
windows and cache-read share remain unknown when the ledger did not record them.
Reconstruct an event after it happened, entirely from local ledgers:
agentguard-burn record --session 0b0c2202 --from 2026-09-20T20:20:32Z --to 2026-09-20T20:20:52Z --out run.jsonl
agentguard-burn render run.jsonl --mp4 run.mp4 --gif run.gif
agentguard-burn render run.jsonl --mp4 run-phone.mp4 --phone
--phone renders a 1080 by 1350 frame with a stacked 40 by 36 cell layout (spawn count, spawn attempts against the ceiling, one block per copy that launched, the last three decisions, the STOP card and an explainer), so a recorded run stays readable on a phone screen. The provenance file records the layout and frame size. --label <text> (1 to 40 printable characters) names what the recording was, for example --label "deliberate fan-out test"; it is drawn in the frame (phone: the explainer row; wide: the header; dashboard: the session line) and recorded in the provenance file. It is never inferred from the recording. --playback <rate> (1 to 4) renders a faster cut of the same recording; the frame says the rate ("RECORDED · 1.5x" or "PLAYBACK 1.5x") and provenance records it, while frame times stay in recorded time.
record writes each decision at its real timestamp, with one frame per second
between decisions. Bounds are optional and otherwise use the session's first
and last decisions. Earlier decisions still supply window context. The output
must be new. It contains identifiers and measured metadata, never tool input,
output or transcript content. See the live panel and September 20 evidence.
agentguard-burn status shows this UTC month’s recorded spawn stops and blocked shell commands. It reads Burn’s local ledger and plugin ledgers registered by either host, joins mirrored spawn receipts once, and excludes shadow decisions, approval holds and outcomes. Missing command history is unknown. Nothing is uploaded. Free installations see one Team line per UTC month. Counts update on later status calls without repeating that line.
Run agentguard-burn quiet on, or node runtime/policy-cli.cjs quiet on from the AgentGuard plugin directory, to dismiss the monthly summary, the why tip, the weekly STOP invitation and plugin version announcements on this machine. Both write the same flag file, upgrade-moments/quiet, and use AGENTGUARD_HOME or the default local AgentGuard home for it. agentguard-burn quiet off removes that file.
FAQs
Local session usage and runaway-agent circuit breaker. Explain tokens, cache rewrites and API list cost, track pace, warn on heavy turns, and gate agent fan-out with signed receipts. Developer command ranking requires explicit --rank.
The npm package @agentguard-run/burn receives a total of 1,747 weekly downloads. As such, @agentguard-run/burn popularity was classified as popular.
We found that @agentguard-run/burn demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Security News
/Company News
Capital One is partnering with Socket to proactively secure its open source supply chain.

Security News
Socket CTO Ahmad Nassri discusses how to keep AI agents from bypassing package blocks, limit credential access, and monitor their actions.

Security News
GPT-6 Astra tried to plant malicious code in simulated open source projects using fake GitHub accounts and deceptive PRs during an assigned CTF challenge.