
Security News
Ruby's Bundler 4.0.18 Extends Cooldown to bundle lock and bundle cache
The supply chain control that delays freshly published gems now covers lockfile generation and gem vendoring in Ruby projects.
hotato
Advanced tools
Find what broke in your agent calls. Pin it so it never ships again. Local call forensics and regression guards for AI agents: turn timing, and whether the agent did what it said. Nothing leaves your machine.
Find what broke in your agent calls. Pin it so it never ships again.
pip install hotato
hotato check --demo # a bundled failing call: what broke, and when
hotato check ./call.wav # then your own recording
hotato vapi health # or your last 100 Vapi calls
Zero config. Vapi, Retell, Bland, Synthflow, Millis, or local audio. Works with any recording: mono, dual-channel, or a transcript. Timing math, not a judge. Offline. Free. MIT.
pip install hotato
export VAPI_API_KEY=...
hotato vapi health --last 7d --output report.html
Open report.html: every critical incident, timestamped, and your Voice Stability Score.
export RETELL_API_KEY=...
hotato retell health --call-id CALL_ID
--call-id is required and repeatable: you name the Retell calls to pull.
hotato bland health, hotato synthflow health, and hotato millis health
follow the Vapi shape. Those stacks mix both voices onto one channel, so
they measure silence timing, dead air and latency gaps, each finding with
its measured confidence; barge-in and talk-over need two channels.
hotato autopsy ./call.wav
Writes a self-contained HTML report to hotato-output/, plus the JSON
pin reads. Open the HTML in your browser.
autopsy turns one recording into timestamped incidents. scan reads a
folder and tracks the trend. When you are ready, move a finding into CI:
hotato pin turns one incident into a portable failure check, and hotato prove re-runs every stored check and fails the build rather than pass on
evidence it cannot re-read. Every verdict carries its own evidence across
five dimensions: outcome, policy, conversation, speech, reliability.
For continuous use: run hotato vapi health on a schedule, and open
hotato console --production-db evidence.db to watch calls land live.
The exit code is the verdict: 0 pass, 1 fail, 2 refuse (could not tell).
# .github/workflows/voice-qa.yml
on: [pull_request]
jobs:
hotato:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: attenlabs/hotato@v1.19.1
with:
contracts: contracts/
hotato-version: 1.19.1
contracts/ holds what pin wrote. Full workflow: docs/CI.md.
Point Claude Code, Cursor, or any coding agent at this repo: it reads
AGENTS.md and runs the loop end to end offline, no key. Over local
stdio the MCP server adds the scorer plus read/verify/propose tools:
uvx --from "hotato[mcp]" hotato-mcp (docs/MCP.md). Deploying is yours.
hotato runs offline, on the machine that invokes it. The core is stdlib-only Python: no account, no key, no network call of its own. Your traces, prompts, and audio stay on your disk. Scoring with a language model you host yourself is a separate opt-in add-on, outside that core.
The whole loop, command by command: docs/LIFECYCLE.md.
First touch to a CI gate: docs/GETTING-STARTED.md.
Feed it what you already have: docs/CONNECT.md ·
docs/TRACE.md · docs/SIMULATE.md.
What every verdict stands on: docs/EVIDENCE-CONTRACT.md.
Next to the hosted alternatives: docs/COMPARE.md.
The deep toolkit -- capture, simulation, load, benchmarking, the fix ladder,
the fleet control plane -- lives under hotato lab (hotato lab --help).
The public commands are durable. hotato lab moves faster, and every
command name that worked before 1.17 still runs unchanged.
| Property | Value |
|---|---|
| Footprint | ~10 MiB installed, 0 runtime dependencies (stdlib-only) |
| Reproducibility | byte-for-byte: the same recording, the same report |
| Exit codes | 0 pass · 1 fail · 2 refuse |
| Release integrity | OIDC Trusted Publishing + build-provenance attested |
| Runtime | offline, off the production data path |
PYTHONPATH=src python3 -m hotato.benchmark \
--scenarios corpus/real/scenarios --audio corpus/real/audio
On 13 recorded AMI Meeting Corpus clips, the median error between measured caller-onset and the human word-alignment label is 20 ms. Provenance: corpus/real/README.md · method: METHODOLOGY.md.
Timing is measurable only when the two voices arrive on separate channels; a mono or mixed export is marked NOT SCORABLE and refused (hotato trust --stereo call.wav). The full four-tier evidence policy (what each verdict stands on, per input) is docs/EVIDENCE-CONTRACT.md.
Issues and PRs welcome: CONTRIBUTING.md · SECURITY.md · CHANGELOG · docs/
MIT (LICENSE)
mcp-name: io.github.attenlabs/hotato
FAQs
Find what broke in your agent calls. Pin it so it never ships again. Local call forensics and regression guards for AI agents: turn timing, and whether the agent did what it said. Nothing leaves your machine.
The pypi package hotato receives a total of 656 weekly downloads. As such, hotato popularity was classified as not popular.
We found that hotato demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.
Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Security News
The supply chain control that delays freshly published gems now covers lockfile generation and gem vendoring in Ruby projects.

Security News
During a UK cyber test, a Mythos 5 agent used sockpuppets, social engineering, and prompt injection to try to get a maintainer to merge malware.

Company News
Socket is now in the AWS Security Hub Extended plan. Adopt it through AWS, apply committed spend, and block malicious open source packages.