
Product
Introducing Socket Scanning for VS Code Marketplace Extensions
Socket now scans VS Code extensions, giving teams early detection of risky behaviors, hidden capabilities, and supply chain threats in developer tools.
@netlify/axis
Advanced tools
Open source tooling and scoring framework to measure how well services work for AI agents.
AXIS is an open source tooling and a scoring framework to measure how well services work for AI agents. Think Lighthouse, but for agent experience.
Give AXIS a scenario, an agent, and a prompt. It runs the agent, captures a full transcript, and produces a graded score across four independent dimensions: Goal Achievement, Environment, Service, and Agent.
The web has Lighthouse. APIs have contract testing. Performance has k6. But there's no standardized way to answer: "How well does my system work when an AI agent tries to use it?".
As agents become a primary interface for interacting with sites, APIs, and developer platforms, the systems they interact with need to be measured and optimized for that experience — just like we optimize for page load time or accessibility. AXIS is that measurement.
npm install @netlify/axis
axis.config.json:
{
"scenarios": "./scenarios",
"agents": ["claude-code"]
}
scenarios/hello-world.json:
{
"name": "Hello world",
"prompt": "Navigate to https://example.com and describe what you see on the page.",
"judge": [
{ "check": "Agent visited the target URL", "weight": 0.5 },
{ "check": "Agent provided a description of the page content", "weight": 0.5 }
]
}
axis run
AXIS executes the scenario, scores the result, and writes a report to .axis/reports/.
Agents are stochastic, so the same scenario can score differently on two identical runs. Set runs above 1 to sample a pair several times:
axis run --runs 3
{
"settings": { "runs": 3 }
}
AXIS keeps every run but headlines a single representative one: the real run whose composite sits nearest the median. Since run counts are odd, that run's score is the median, so the headline equals the median of the runs listed beneath it and the live terminal output matches the report exactly. It also keeps the score explainable, since the transcript and audits you drill into belong to the run being reported, which an average would not.
Alongside it the report records the spread (median, range, sigma), each run's four dimension scores, and a reliability fraction. A crashed run is excluded from the score and counted against reliability instead, so flakiness never masquerades as low quality; a run whose score was withheld because judging failed leaves the denominator entirely rather than being charged to the agent.
The payoff lands in --compare-baseline: a baseline built from repeats knows its own standard deviation, so a regression is a move that exceeds the noise the suite actually measured rather than a fixed 1-point threshold.
Run counts must be odd (1, 3, 5, up to 19). An even sample has no middle run, so the median falls between two runs and the headline belongs to none of them; at runs: 2 the selection degenerates entirely and always returns run 1 regardless of merit. Even values are rejected rather than rounded.
Repeats default to off. Each extra run is a full agent execution plus its judge calls, so cost scales linearly. See running tests for scheduling, AXIS_RUN_INDEX, and the report layout.
Full documentation lives at axis.run:
axis.config.json, scenarios, MCP servers, skillsaxis run, axis reports, axis baselineUse the programmatic API when you want to integrate AXIS into an existing test runner, build tool, or CI pipeline rather than calling the CLI directly.
run({ profile }) and loadConfig(path, { profile }) select a named profile from the config, the same overlay axis run --profile applies. loadConfig returns the resolved config alongside the pre-merge baseConfig, for callers that need the suite layout as a whole rather than the active suite.
Delivered: scenario runner, four-dimension scoring pipeline, baselines with regression detection, repeated runs with representative-run selection and noise-aware regression bands, MCP/skills wiring, custom adapter API, config profiles for running one repo's scenarios under several agent matrices, built-in adapters for Claude Code, Codex, and Gemini.
Planned:
AXIS is built in the open. Contributions are welcome. New scenarios, agent adapters,
bug fixes, and documentation improvements all help.
AXIS is open source under the MIT license, created by Netlify and developed with founding contributors including Auth0 and Resend.
Full docs: axis.run
FAQs
Open source tooling and scoring framework to measure how well services work for AI agents.
The npm package @netlify/axis receives a total of 1,453 weekly downloads. As such, @netlify/axis popularity was classified as popular.
We found that @netlify/axis demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 9 open source maintainers collaborating on the project.

Product
Socket now scans VS Code extensions, giving teams early detection of risky behaviors, hidden capabilities, and supply chain threats in developer tools.

Research
/Security News
Socket uncovered two malicious VS Code themes in a GlassWorm-linked cluster with thousands of installs across VS Code Marketplace and Open VSX.

Security News
/Company News
Capital One is partnering with Socket to proactively secure its open source supply chain.