
Product
PHP and Composer Support Is Now in Beta
Socket’s PHP and Composer support is now in Beta for all customers, with PHP reachability analysis generally available.
evolveguard-cli
Advanced tools
Regression-testing CI gate for self-edited Claude Agent Skills (SKILL.md, MEMORY.md): golden-transcript record/replay against a skill's own declared and inferred capability surface, zero hosted infrastructure.
Regression-testing CI gate for self-edited Claude Agent Skills -- SKILL.md
manifests and Claude Code auto-memory MEMORY.md files -- catching
behavioral drift before an edit ships.
Claude Code's Agent Skills can be authored by a human, or by an agent
itself: /skillify turns a workflow into a SKILL.md, and the auto-memory
system in this same environment writes MEMORY.md files that quietly
change what an agent does in its next session. None of that gets a
regression check by default. A skill edit that breaks a working workflow
looks exactly like a skill edit that fixes one, until someone notices the
agent stopped doing something it used to do. evolveguard records a
baseline of a skill's own capability surface (what tools it's declared or
shown to use), then re-derives that surface every time the skill file
changes and diffs the result against the baseline: same tool-call
sequence, or a flagged drift with a specific reason.
pip install evolveguard-cli
or with uv:
uv add evolveguard-cli
Live on PyPI at pypi.org/project/evolveguard-cli. The package was originally published under the plain name
evolveguard; that older PyPI project is retired and no longer receives updates -- installevolveguard-cli(above) instead. To install from source anyway:git clone https://github.com/RudrenduPaul/evolveguard.git cd evolveguard/python && pip install -e .
About the npm package: evolveguard-cli is also live on npm at
npmjs.com/package/evolveguard-cli
-- the TypeScript/JavaScript CLI and library, same record/replay/diff
pipeline, same CLI flags. Renamed 2026-07-19 from the old plain
evolveguard, which is now deprecated. Install it with
npm install -g evolveguard-cli, or clone the repo and run npm run build && npm link if you need it from source.
# 1. Record a baseline against a skill and its labeled fixtures
evolveguard record ./skills/my-skill/SKILL.md --fixtures ./fixtures/my-skill.json
# writes ./skills/my-skill/.evolveguard-baseline.json
# 2. Edit the skill (by hand, or let an agent edit it)
# 3. Check for drift
evolveguard check ./skills/my-skill/SKILL.md
# writes ./evolveguard-report.json, exits 1 if drift was found
A fixtures file is a JSON array of labeled prompts and the tool-call shapes each one is expected to touch:
[
{
"id": "scan-a-monorepo",
"prompt": "scan a monorepo",
"expectedToolCalls": [{ "tool": "fs.read" }, { "tool": "fs.write" }]
}
]
expectedToolCalls is optional -- omit it and the fixture is treated as
exercising the skill's entire capability surface. scopeMatches (a glob)
narrows a tool to a specific filesystem scope, e.g.
{ "tool": "fs.write", "scopeMatches": "./workspace/**" }.
Or call the library directly (the agent-native path):
from evolveguard import record_baseline, replay_skill, diff_all, write_baseline, read_baseline
baseline = record_baseline("./SKILL.md", "./fixtures.json")
write_baseline("./.evolveguard-baseline.json", baseline)
# ... skill gets edited ...
saved = read_baseline("./.evolveguard-baseline.json")
replay = replay_skill("./SKILL.md", saved)
report = diff_all(saved, replay)
print(report.summary, report.exit_code)
evolveguard does not run a live LLM agent, and it does not replay a real conversation transcript. It is a static, deterministic tool by design:
record parses a skill file's YAML frontmatter (declared tools,
network, filesystem, scope, and any bundled hooks), scans the
skill's body text and any hook scripts for static evidence of network
calls or filesystem writes, and combines both into a capability
surface -- the set of tools the skill is declared or shown to use.
Each fixture's expectedToolCalls filters that surface down to the
tools the fixture author says it cares about; the result is the
recorded baseline.check re-reads the (possibly edited) skill file, re-derives its
capability surface with the exact same logic, and re-filters it per
fixture.diff compares baseline vs. current per fixture (PASS if the tool
set and scopes match, DRIFT with a specific reason otherwise), and
separately diffs the whole capability surface so a new capability that
no fixture's expectedToolCalls happened to cover still gets caught.evolveguard detects changes in what a skill is declared or shown to be
capable of. It can't tell you whether a live agent run would actually
behave differently on a given prompt -- that's a real, intentional scope
limit. The tradeoff: it just needs a SKILL.md file and a fixtures file,
with nothing hosted and no SDK to integrate against, which is also why it
runs fully offline in a pre-commit hook or CI job. This Python package is
a genuine, independent port of the pipeline -- not a wrapper around the
Node binary. See the
project README for
the fuller design writeup.
Braintrust is a general LLM eval and observability platform -- a
strong choice if you're already logging traces from a live agent and want
statistical eval scoring across runs, but it needs SDK integration and an
eval-definition step per app. evolveguard needs neither: point it at one
SKILL.md file and a fixtures JSON, and it works, with no live LLM calls.
agent-eval (this same
author's other repo) answers a different, more general question: "did my
agent's behavior change between two versions I define," for any agent,
framework-agnostic, requiring you to stand up and run both versions
yourself. evolveguard answers a narrower question triggered directly by a
file diff on SKILL.md/MEMORY.md: did this specific edit change the
capability surface the baseline recorded, with nothing to stand up or run.
See the
project README's comparison table
for the full breakdown.
usage: evolveguard [-h] [-V] {record,check,report,mcp} ...
Regression-testing CLI for self-edited Claude Agent Skills (SKILL.md,
MEMORY.md) -- golden-transcript record/replay against a skill's own
declared and inferred capability surface, zero hosted infrastructure.
positional arguments:
{record,check,report,mcp}
record Record a golden-transcript baseline for a skill
against a set of labeled fixtures
check Replay the fixtures from a baseline against the
current (possibly edited) skill and report drift
report Print a previously generated evolveguard-report.json
mcp [coming soon] Expose record/check/report as MCP
tools for a coding agent to call mid-session
options:
-h, --help show this help message and exit
-V, --version show program's version number and exit
evolveguard record <skillPath> --fixtures <path> [--baseline <path>] [--json],
evolveguard check <skillPath> [--baseline <path>] [--report <path>] [--allow-drift] [--json],
and evolveguard report [reportPath] [--json] mirror the npm CLI's flags
and defaults exactly -- see the
project README's CLI reference
for the full --help output of each subcommand.
Exit codes: 0 all fixtures PASS and no surface-level drift, 1 at
least one DRIFT was found (pass --allow-drift to still exit 0 while
still reporting it), 2 a usage error or a file that failed to parse.
Every subcommand supports --json for structured output an agent can
parse directly:
evolveguard check ./SKILL.md --json
evolveguard mcp (the argparse subcommand) is documented but not
implemented as its own MCP server -- it now delegates to the real MCP
server described below. The npm/TypeScript distribution's evolveguard mcp is still a "coming soon" stub; call record/check/report --json
directly as a subprocess (or the library functions in-process) from your
coding agent if you're on that distribution.
This package ships a Model Context Protocol server, so an MCP-compatible agent (Claude Desktop, Claude Code, etc.) can call evolveguard directly instead of shelling out and parsing text.
pip install "evolveguard-cli[mcp]"
Add it to your MCP client's config, for example Claude Desktop's
claude_desktop_config.json:
{
"mcpServers": {
"evolveguard": {
"command": "evolveguard-mcp"
}
}
}
It exposes a single tool, run(args: list[str]), that shells out to the
installed evolveguard CLI with the exact argv you'd type at a terminal
and returns {returncode, stdout, stderr, json?} (or {error: ...} if the
command fails, times out, or exits non-zero) -- one tool covers record,
check, and report without a bespoke MCP tool per subcommand. Example
call from an agent:
{ "tool": "run", "arguments": { "args": ["check", "./SKILL.md", "--json"] } }
which returns the same structured report evolveguard check ./SKILL.md --json would print, plus the raw returncode/stdout/stderr. You can
also run evolveguard mcp (the CLI subcommand) as an equivalent to
evolveguard-mcp if the mcp extra is installed.
No accuracy claim ships without the command that produced it. From a clone of the repo:
cd python && pytest tests/test_benchmark.py -v
against fixtures/labeled-non-breaking-edits/ (shared with the
TypeScript test suite, not duplicated) -- a small, hand-labeled corpus of
5 real before/after SKILL.md pairs: 2 labeled non-breaking (a wording
tweak, a typo fix) and 3 labeled breaking (a filesystem-scope widen, a new
write capability, and a hook script gaining a network call). As of this
release: 0% false positives (0 of 2 non-breaking cases flagged as
drift), matching the npm package's own documented result on the same
corpus. The corpus is small and will grow as more real skill edits are
reported.
What does this package do?
It detects capability drift in Claude Agent Skill files (SKILL.md) and
Claude Code auto-memory files (MEMORY.md) after they're edited, by a
human or an agent. It records a baseline of what a skill is declared or
shown to use, then re-derives that surface after an edit and diffs the
two, flagging drift with a specific reason instead of letting a broken
edit ship silently.
How does this differ from the npm package?
Nothing behaviorally -- evolveguard-cli on PyPI is a genuine, independent
Python port of the same TypeScript pipeline on npm (also evolveguard-cli),
not a wrapper around the Node binary. Both parse the same skill-file
schema, produce the same capability surface, and share baseline/report
JSON files interchangeably. Pick whichever language fits your existing
toolchain.
Is it safe to run against an untrusted or self-edited skill file?
Yes. evolveguard reads local files and never executes any of them -- no
eval, no subprocess, no dynamic import of scan-target content. See
"Security" below and SECURITY.md
for the full scope notes.
Does it need API keys or make network calls?
No. record and check are both fully static and deterministic -- no LLM
calls, no hosted service, no network access at any point in the pipeline.
That's also why it runs fully offline in a pre-commit hook or CI job.
Does it work with MEMORY.md files, which have no frontmatter?
Yes. A file with no frontmatter is parsed with an empty declared scope, so
its capability surface comes entirely from static evidence found in the
body text.
See CONTRIBUTING.md.
There is no enforced minimum coverage threshold; the bar is that the full
pytest suite (pytest from python/) passes and new behavior ships with
tests.
cd python
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest
evolveguard reads local files you point it at and never executes any of
them -- no eval, no subprocess, no dynamic import of scan-target
content, no network calls. Hook script paths declared in a skill's
frontmatter are resolved and validated against that skill's own directory
before being read, including a symlink-escape check, so a malicious or
broken skill file cannot make evolveguard read outside its own folder.
Baselines and reports are read and written as plain JSON, never pickled or
otherwise deserialized as executable data. See
SECURITY.md
for the disclosure process. Honest note: this project does not
currently publish SLSA provenance, Sigstore signatures, or an SBOM, and
has no OpenSSF Scorecard badge set up -- none of that infrastructure
exists yet for either distribution, so it isn't claimed here.
MIT, see LICENSE.
FAQs
Regression-testing CI gate for self-edited Claude Agent Skills (SKILL.md, MEMORY.md): golden-transcript record/replay against a skill's own declared and inferred capability surface, zero hosted infrastructure.
We found that evolveguard-cli demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 2 open source maintainers collaborating on the project.

Product
Socket’s PHP and Composer support is now in Beta for all customers, with PHP reachability analysis generally available.

Product
Socket is bringing experimental protection to Firefox, scanning 97,000+ extensions in Mozilla's official directory for malware and risky updates.

Research
/Security News
Three compromised Rust crates pulled in a malicious dependency that downloaded and executed cross-platform malware during Cargo builds.