New:Microsoft Teams Notifications Are Now Available in Socket.Learn more
Get Started

codecalc

Package Overview
Dependencies
Maintainers
1
Versions
14
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

codecalc

Universal code + logic calculator for AI models: execute 31 languages, evaluate logic, analyze complexity. Exposed via MCP.

Source
pipPyPI
Version
0.13.0
Weekly downloads
951
Maintainers
1
Created

codecalc — universal code & logic calculator for AI models

codecalc is an offline, self-hosted MCP server that gives an AI agent a calculator, a code runner, and a logic checker — so it gets a correct answer instead of a guessed one. It runs code in 31 languages, does exact symbolic math, solves SMT/logic problems, and measures complexity, all exposed as 51 MCP tools.

Fastest path: uvx 'codecalc[full]' setup --write registers codecalc with your MCP client automatically. New to MCP, or want more detail first? See QUICKSTART.md, or the Install section below.

Three things nobody else offers together cleanly:

  • Offline-core — ships no model, no API key, no gateway, no telemetry. The core opens no sockets; network access is opt-in and only where a specific tool's job needs it (the Piston provider, install_package, the runtime-update tools, executed code unless no_net, and a one-time in-process grammar download on first analyze_complexity — full breakdown in the network-boundary table below).
  • Safe execution of untrusted code — an opt-in strict isolation boundary (gVisor+Docker on Linux, AppContainer on Windows) layered above the default rlimit sandbox, fail-closed and attested.
  • Verification toolsverify_translation proves a port to another language behaves identically, verify_optimization proves an optimization preserved behavior, and z3_check proves or refutes logic with an SMT solver.

When to use codecalc

Use it when you want a free, local, private, hardened code-runner and verifier that an MCP agent can call directly — no vendor account, no cloud spend, nothing leaving the machine except where a tool's job explicitly requires it.

Reach for something else when you want managed cloud scale instead of self-hosting (a hosted sandbox like E2B or Modal), or when you're not self-hosting at all and the model vendor's built-in code interpreter already covers what you need.

codecalc vs. the alternatives

codecalc is not a general cloud sandbox and not a vendor code interpreter. It overlaps with several things and beats them in only one narrow place — forcing a model to measure a claim instead of asserting it. Where that isn't what you need, one of these is the better tool, and this table says so plainly.

You want…Better fitWhy
To just run some Python/JS quickly, zero setupYour model vendor's built-in interpreterAlready there, already sandboxed, nothing to install. Anthropic's code-execution tool has internet access "completely disabled" and cannot install packages at runtime; OpenAI's hosted containers have no outbound network access by default, with an org-level network_policy allowlist as an opt-in. Both return output artifacts by reference (Anthropic a file_id via the Files API, OpenAI a container_file_citation) rather than inline (Anthropic code-execution tool docs, https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool; OpenAI shell/container tool guide, https://developers.openai.com/api/docs/guides/tools-shell; both retrieved 2026-09-07)
Heavy or multi-tenant workloads, managed scaleA cloud sandbox (E2B, Modal, Daytona)Per-tenant Firecracker/gVisor isolation codecalc does not claim by default
Pure arithmetic or symbolic math, nothing elseA small calculator or SymPy MCPLower token cost; none of the 31-language runtime machinery
A model to stop guessing numbers, equivalence, and speedups — locally, privately, with graded evidencecodecalcExact rationals, verify_translation/verify_optimization, and unenforced/grade honesty — offline, no account

Do not reach for codecalc if you need multi-tenant or network-exposed isolation (its threat model is explicitly single-operator, local, stdio), if zero-setup convenience matters more than measurement, or if a hosted interpreter already covers your case. It earns its keep only when the correctness of the claim — not just "it ran" — is the point.

Install

Quickstart with codecalc setup

The fastest path to a working MCP connection, without reading the rest of this section:

uvx 'codecalc[full]' setup            # prints what it would do — nothing on disk changes
uvx 'codecalc[full]' setup --write    # applies it: merges your client's config, copies the skill

It detects which MCP client is installed (Claude Desktop, Claude Code, Cursor, VS Code, Zed — pass --client=NAME if none or several are found), reuses codecalc doctor's own backend/extras/grammar-cache checks, prints the exact config block in that client's own JSON shape with absolute paths already filled in, runs two real canaries (execute_code, evaluate_expression) to prove the connection would work, and ends in one verdict: ready / degraded / not-ready. --write is the only mode that changes anything — it MERGES the codecalc entry into your existing client config (every other server stays exactly as it was) and backs up the original to <path>.codecalc-bak first. codecalc --help lists every subcommand.

[!NOTE] Published as codecalc 0.13.0 on PyPI (pip install codecalc) and the codecalc-exec 0.13.0 executor on crates.io (#91). Every release artifact carries a keyless sigstore build-provenance attestation — verify one with gh attestation verify <file> --repo The-40-Thieves/codecalc; PyPI wheels additionally carry PEP 740 attestations.

Where to find codecalc

WhereWhat you getLink
PyPIpip install codecalc / uvx codecalcpypi.org/project/codecalc
crates.iothe codecalc-exec Rust executor cratecrates.io/crates/codecalc-exec
GitHub Releaseswheels for every platform, the executor binaries, the .mcpb bundle, an SBOM, and SHA256SUMSgithub.com/The-40-Thieves/codecalc/releases
MCP registry (official)the io.github.The-40-Thieves/codecalc server entry server.json publishes toregistry.modelcontextprotocol.io
Smitheryhosted listing and one-click client configsmithery.ai/servers/@The-40-Thieves/codecalc
Glamahosted listing and the score badge aboveglama.ai/mcp/servers/The-40-Thieves/codecalc
MCPB (Claude Desktop)the drag-and-drop bundle, attached to every GitHub Releasesee GitHub Releases, above
Docker MCP Catalogthe mcp/codecalc image Docker builds from this repo's docker/mcp-server.Dockerfile, for Docker Desktop's MCP Toolkit (docker mcp server enable codecalc)submitted as docker/mcp-registry#5025; listed at hub.docker.com/mcp/server/codecalc once merged

Not yet listed: PulseMCP and mcp.so do not carry a codecalc entry yet; the Docker MCP Catalog entry is pending review. PulseMCP's own submission page (checked 2026-09-08) says it is not accepting new submissions and that publishing to the official MCP registry — already done, row above — is what it indexes from once submissions reopen, so there is nothing to submit there today. mcp.so takes a submission through its own form. See docs/distribution.md for the exact steps, kept there rather than here because submitting is an action for whoever runs it, not a fact about the current release.

The published install (simplest — no build step, and what most people want):

uvx 'codecalc[full]'          # run it directly, no environment to manage
# or
pip install 'codecalc[full]'  # into your own virtualenv

From source, if you would rather build the executor yourself:

git clone https://github.com/The-40-Thieves/codecalc
cd codecalc
uv sync --all-extras                 # or: pip install -e '.[full]'
cargo build --release --manifest-path executor/Cargo.toml
mkdir -p bin                         # bin/ is gitignored, so a fresh clone has none
cp executor/target/release/codecalc-exec bin/
uv run codecalc doctor               # verify: backend should read `rust`

Without the cargo build, everything still runs on the pure-Python fallback — doctor will say so, and the network table below says what that costs.

Why [full]. The base install is the MCP surface and the sandbox executor: 31 language runtimes, sessions, packages, ~32 MB. The symbolic half — sympy and z3 — is 88.6 MB measured, and a caller who only runs code should not download an SMT solver to do it. So it is an extra:

installsizewhat you get
codecalc~32 MBexecute_code, sessions, packages, complexity-free tools
codecalc[symbolic]+83 MBevaluate_expression, solve, limits, truth tables, z3, units
codecalc[parsing]+5 MB installed, +89 MB fetched on first useanalyze_complexity via tree-sitter
codecalc[full]~120 MBeverything

Nothing fails silently: a tool whose extra is missing returns {"ok": false, "error": "sympy is not installed. It ships in the 'symbolic' extra: pip install 'codecalc[symbolic]' ..."}, and codecalc doctor lists which extras are present before you make a call.

Editions

Four names for the capability sets above, plus the two that live outside pyproject.toml entirely — a Docker image and an opt-in isolation boundary. The invariant that makes "edition" a meaningful word here: in the edition that lists a tool, that tool is functional — never listed-but-missing-its-extra. A tool an edition doesn't have returns the dependency_missing contract error naming the extra that provides it (see above), not a silent failure or a tool that appears to exist and doesn't work.

EditionInstallWhat you get
Fulluvx 'codecalc[full]' / pip install 'codecalc[full]'The recommended local product: the native (Rust) executor, symbolic tools (evaluate_expression, symbolic, z3_check, …), and parsing (analyze_complexity). Everything this README documents actually runs.
Coreuvx codecalc / pip install codecalcExecution + non-symbolic tools only — the base install in the table above. Every symbolic/parsing tool is still listed by tools/list (MCP doesn't support per-install schemas), but calling one returns dependency_missing naming the extra, before any other work happens.
Dockerdocker build -f docker/mcp-server.Dockerfile .The MCP server itself, packaged to run as an ordinary container. Core-shaped by default: ships python3/node/ruby/php/perl/gawk/lua/c/cpp/jq/sqlite3 and the default rlimit sandbox — symbolic/parsing are absent by design (no [full] in the base image; see the Dockerfile's own comment for why, including an arm64 z3-solver wheel gap). --build-arg CODECALC_EXTRA=full adds them. This image cannot nest the Strict Host boundary below inside itself (no privileged docker-in-docker), and codecalc doctor inside it says so rather than claiming a boundary it doesn't have.
Strict Hostopt-in; CODECALC_STRICT_URL (client) or the gVisor+Docker host itself (server) — see docs/deployment/README.mdNot an install, a boundary: the gVisor runsc sandbox on Linux, or AppContainer hardening on Windows, layered above whichever install above is already running. Fails closed — no digest pinned, no fallback to unenforced local execution.

codecalc doctor reports which of these you're actually running (backend, extras present, strict_runtime prerequisites) — read it before assuming a capability rather than after a tool call surprises you.

.github/workflows/release.yml publishes a platform-tagged wheel per target (Linux x86_64/aarch64 musl, macOS x86_64/aarch64, Windows x86_64), each carrying the matching codecalc-exec binary and — where the platform has one — its --no-net shim, so executor.backend() == "rust" on install without a manual build step. No wheel for your platform, or installed from source instead? Everything still runs; see the network table above for what falls back and to unenforced in that case.

Point an MCP client at the installed command. The key differs by clientmcpServers for most, servers for VS Code, context_servers for Zed — so these are given separately rather than as one snippet to adapt:

Claude Desktop~/Library/Application Support/Claude/claude_desktop_config.json (macOS), %APPDATA%\Claude\claude_desktop_config.json (Windows) · Cursor (.cursor/mcp.json) and Claude Code (.mcp.json) use the same shape:

{ "mcpServers": { "codecalc": { "command": "uvx", "args": ["codecalc[full]"] } } }

VS Code.vscode/mcp.json, top-level key is servers:

{ "servers": { "codecalc": { "command": "uvx", "args": ["codecalc[full]"] } } }

Zed~/.config/zed/settings.json, key is context_servers:

{ "context_servers": { "codecalc": { "command": "uvx", "args": ["codecalc[full]"], "env": {} } } }

Windows paths need doubled backslashes in JSON. If you installed into a venv rather than using uvx, point at the interpreter directly:

{ "mcpServers": { "codecalc": {
    "command": "C:\\path\\to\\venv\\Scripts\\python.exe",
    "args": ["-m", "codecalc"] } } }

Run codecalc doctor to print a config block with the absolute paths of your install already filled in.

Install the skill too. The tools cannot help a model that never reaches for them — a model confident about 0.1 + 0.2 does not feel uncertain, it feels finished. codecalc/SKILL.md ships inside the package and says when calling is mandatory (any non-integer, any comparison you will state, anything past 2^53, any number stated as a claim), when it is noise (2 + 3 + 4 needs no tool), and how results must be reported — passed: true means "equivalent on N inputs", never "verified". codecalc doctor prints its path; copy it into your client's skills directory. check_claims.py gates it, so it cannot name a tool that does not exist or a field no tool returns.

Not sure what your install actually resolved? Ask it, rather than finding out from a tool call later:

codecalc doctor          # or: python -m codecalc doctor

This is the install verification step. It exits 0 when the install can execute — a writable workspace and a resolved backend — and 1 when it cannot, so it works unchanged in a Dockerfile, a provisioning script or a CI job. A missing optional extra or an uninstalled Haskell does not fail it: those are facts about the host, not a broken install, and a check that goes red for them is one people learn to ignore.

It prints the execution backend and the binary behind it, whether installs are confined, the status of every one of the 31 runtimes, whether the workspace is writable, and a client config block with absolute paths filled in. All of that is otherwise discoverable only by making a tool call and reading backend, unenforced, or a failure.

codecalc doctor --json   # the same report, for scripts
codecalc doctor --deep   # actually RUN each runtime, and read its version

--json emits the report and nothing else, against a published schema (docs/contract/doctor-v1.schema.json) carrying the same contract_version and the same policy as a tool result.

Each runtime reports one of four states, and the difference between two of them is which measurement was actually taken:

statemeans
supportedcodecalc knows the language; nothing for it resolves here
installedits command resolves and is executable — not run
unhealthyresolves but cannot run, or was run and failed
availableactually executed here and answered — --deep only

status_basis says which pass produced them. Without --deep nothing is ever reported available, because nothing was executed, and claiming otherwise for a binary that was merely found on PATH would be a stronger measurement than was taken.

Under --deep, a runtime whose version probe never gets an answer (a spawn failure or a timeout) is unhealthy too; a nonzero exit alone only counts when the flag used is one confirmed correct for that command (go version, lua -v, zig version — none of them speak GNU --version, so a bare nonzero exit there is reported as merely unmeasured, not broken). A captured failure lands in probe_error, never in version, which holds a version string or nothing. A compile-then-run language whose run step needs a SECOND, different tool (kotlin: kotlinc compiles, but run launches java directly) reports installed only when BOTH resolve; detail names whichever half is missing.

Building the Rust core yourself, or running from a checkout? See "Build the Rust core" and "Run the server" below.

Use it from an MCP client

One-click install: both buttons register uvx codecalc[full] (the recommended Full edition) and require uv to be installed.

Add to Cursor Install in VS Code codecalc MCP server

The shortest version of the config above — this registers codecalc as a stdio MCP server. The console entry point is codecalc, so uvx codecalc launches it directly:

{
  "mcpServers": {
    "codecalc": { "command": "uvx", "args": ["codecalc"] }
  }
}

Installed with pip install codecalc instead? Point at the resolved command with no args:

{
  "mcpServers": {
    "codecalc": { "command": "codecalc" }
  }
}

Network boundary

CodeCalc's core opens no sockets. No model gateway or telemetry is built in. tests/test_offline.py asserts this for the top-level core modules. The opt-in Piston provider is the deliberate exception: its wire client lives under codecalc/provider_adapters/ and is registered only when CODECALC_PISTON_URL is configured.

That is a claim about the package, not about every tool call, and the difference is worth stating rather than leaving a reader to discover:

layerreaches the network?
CodeCalc coreNo HTTP client, model gateway, or telemetry. One dependency exception: analyze_complexity may download a grammar on first use (see below)
configured Piston providerYes, explicitly. Calls only the operator-supplied CODECALC_PISTON_URL; credentials stay in its authorization header and are redacted from results
install_packageYes, by design. It runs uv / npm / gem / cargo, which fetch from their registries. Installer hooks also run outside the sandbox — see SECURITY.md
runtimes_status, update_runtimesYes. They shell out to mise / rustup / swiftly / npm, which check remote versions
code you executeYes, unless no_net=True — and that guarantee needs the native executor (seccomp-bpf where the Linux kernel supports it, a symbol shim otherwise; see the guarantee table below), so the pure-Python fallback reports it in unenforced instead of applying it. Set CODECALC_REQUIRE_NATIVE=1 to turn "fallback in use" into a startup failure instead of a result you have to notice by reading unenforced
execute_code / session_run / execute_code_stream / run_submit with declared dependenciesYes, before the sandboxed step, through the confined install_package path. A PEP 723 block (python3) or the dependencies argument is installed BEFORE the code runs — never inside the sandbox — and refused (capability_not_requested, no fetch attempted) when no_net=True was requested or the capability policy denies or strictly limits network. run_submit's install runs on its own background worker, same as the code that follows it — the call itself still returns a run_id immediately

These distinctions are stated precisely on purpose: a guarantee described more broadly than it is enforced is exactly the failure mode this project works to avoid, so "offline-core" is scoped to what the structural test can actually support rather than claimed as a blanket "no network calls".

A PEP 723 block alone, with no dependencies argument, can trigger the install above. execute_code/session_run read the block out of the source text itself — a caller who passes no dependencies argument at all still gets a confined uv/npm subprocess and real egress if the code they submit happens to carry a # /// script block, whenever the refusal rule above does not apply. This is logged distinctly (dependency_install_implicit in the audit trail, alongside install_denied) so an operator can tell "source text alone triggered this" from an explicit install_package/dependencies= call. To disable it: no_net=True on the call, or a deny-network/strict CODECALC_CAPABILITY_POLICY — either one refuses before any fetch, block or no block. execute_code_stream and run_submit read the block the same way execute_code does. compare_execution is the one holdout: it fans out across several languages with no per-language install plumbing behind it, so it REJECTS an explicit dependencies argument with a validation error rather than approximating one, and DISCLOSES rather than silently drops an inline PEP 723 block it finds in a snippet — that row's result carries dependencies: {"status": "unsupported", "reason": ...} instead of installing from it.

Two ceilings govern a dependency-bearing run, not one. The run's own timeout bounds the sandboxed step; it says nothing about installing dependencies FIRST, outside the sandbox. A separate, fixed budget (codecalc.dependencies.DEFAULT_DEPENDENCY_INSTALL_BUDGET_SECONDS, 120s, aggregate across every dependency of one run) bounds that step instead — exceeding it refuses the run with a stamped timeout naming the budget, before the run's own timeout clock even starts. A sessionless run's dependency workdir is also held to a disk quota — reusing CODECALC_SESSION_DISK_QUOTA_MB (below), the same cap a session workspace already has — and a run that grows past it after a successful install is refused with a stamped resource_exhausted naming the measured size and the cap.

The grammar download, stated plainly, because it is the one that is easy to miss. The other three paths above go through a CHILD PROCESS, which is what tests/test_offline.py says it cannot see. This one does not: tree-sitter-language-pack ships a ~5 MB extension and fetches each grammar on first use, in-process, into a local cache — 28 grammars, 89 MB, about 15 seconds on a cold cache. So the first analyze_complexity call for a given language opens a socket from inside the server.

It is verified (the pack checks a signature and raises on a checksum mismatch), it is cached, and it never happens again for that language. So the offline-core claim is scoped to steady state: this first-use grammar fetch is the one in-process exception, which is why it is called out here rather than glossed over.

For an offline or egress-restricted install, warm the cache first — it is one command, and afterwards nothing here reaches the network. If you installed codecalc (pip install/uvx, not a source checkout), scripts/ did not come with it, so use the shipped console script instead:

codecalc-prefetch-grammars                    # installed: fetch all 28 grammars
codecalc-prefetch-grammars --print-cache-dir  # installed: the directory to copy

Building from source? The script still works and calls the same code:

python scripts/prefetch_grammars.py                    # fetch all 28 grammars
python scripts/prefetch_grammars.py --print-cache-dir  # the directory to copy

codecalc doctor reports whether that cache is populated, so this is discoverable before it matters rather than after a tool call degrades.

Architecture (language-per-strength)

LayerLanguageWhy
Executor core (executor/)RustSandbox + rlimits + process-group kill + JSON CLI. No eval() anywhere near user input; memory-safe host; single static binary
Logic layer (codecalc/logic.py)Pythonsympy (symbolic math, equation solving) and z3 (SMT) have no Rust equivalents
MCP server (codecalc/server.py)Pythonthe official mcp SDK (2.0) generates tool schemas from type hints; protocol 2026-07-28

Python orchestrates; Rust executes; sympy/z3 reason. Each layer does what it's best at. The Rust binary is preferred automatically; a pure-Python executor is the fallback if the binary is missing.

Older-computer support

  • No modern instruction-set requirements — rustc targets a generic CPU by default and nothing overrides it. (executor/.cargo/config.toml explains why -C target-cpu=generic is deliberately NOT written there: it would be a no-op that reads like a guarantee.)
  • Static musl builds run on any Linux regardless of glibc version: bin/codecalc-exec-x86_64-musl, bin/codecalc-exec-aarch64-musl (~430K each; the exact size moves with every toolchain bump, so it is not pinned here)
  • Size-optimized profile (opt-level="z", LTO, panic=abort, stripped) — measured, not assumed: against an otherwise identical opt-level=3 build, z came out 1.02 ± 0.26 times faster on the executor's own path (i.e. no detectable difference) while being 16% smaller. The executor spends its time in syscalls, not arithmetic, so there was nothing for a higher optimisation level to speed up.
  • Lazy sympy/z3 imports. Both are imported on first use, so a session that only executes code never pays for them. This claimed "~40ms, not ~600ms" for a long time while being wrong in both directions: the server took 1.9s to start, and sympy was not actually lazy — units.py imported it at module scope and server.py imports units, so every start paid 437ms for it. Deferring that took spawn-to-first-response from 1888ms to 1243ms (measured, median of 7). The remaining ~870ms is the mcp SDK's own import, which is not ours to remove.
  • The fork-bomb measurement is taken once, and only when it is needed. Sizing RLIMIT_NPROC means reading /proc/<pid>/status for every process on the machine. That walk used to run during argument parsing and again for every step: a C compile-and-run opened 1767 status files on a 590-process box to answer one question three times, and --lang notalanguage paid the full cost to produce a one-line error. Measured lazily and cached, an error costs 1.1ms instead of 13.3ms and a compiled run 78ms instead of 104ms.
  • list_languages probes runtime availability and reports which languages actually work on the machine (graceful degradation on minimal installs)

Build the Rust core

cd executor
cargo build --release                          # native
cargo zigbuild --release --target x86_64-unknown-linux-musl   # static x86_64 (uses zig)
cargo zigbuild --release --target aarch64-unknown-linux-musl  # static arm64
# Copy the executable AND its --no-net shim together. build.rs rebuilds the
# shim whenever blocknet.c changes, but the executor looks for it beside the
# BINARY, so installing only the binary leaves the previous shim in place — and
# a stale shim silently enforces the old policy while every "is it there?"
# check still passes. Copy both or neither.
mkdir -p ../bin                                # bin/ is gitignored, so a fresh clone has none
cp target/release/codecalc-exec target/release/blocknet.so ../bin/

Requires: Rust 1.97+, a C compiler for the --no-net shim (the build warns and carries on without one; on macOS, or a Linux kernel without seccomp support, --no-net then reports itself in unenforced rather than pretending — a Linux kernel with seccomp support enforces it in-kernel either way), and cargo-zigbuild for the static cross-builds (zig is used as the linker; no x86_64 GCC needed).

MCP tools (51) + MCP resources

Every session file is also exposed as an MCP resource: codecalc://session/<session_id>/files/<path> — images render inline for the model, text returns as text, other files download.

Graphical results in MCP Apps hosts: verify_translation and verify_optimization also carry an MCP Apps ui:// view (ui://verify-translation/view.html, ui://verify-optimization/view.html) — a per-case diff table and a per-size timing chart, respectively, rendered inline by a host that supports the extension. Both are self-contained (inline CSS/JS, no network, no external assets); a host without MCP Apps support sees exactly today's text/structured result, unchanged.

Exact arithmetic & programmer-mode: exact rationals, threshold checks, bit analysis, binary64 introspection.

ToolDescription
calc_exactEXACT arithmetic: 0.1+0.2 == 0.3 is True; arbitrary-precision ints, bitwise ops inline, whitelisted math funcs, pi/e/tau
compare_thresholdExact threshold verdict with shortfall: ('1/25', '>', '0.05') → False, shortfall 1/100
percentageExact share and percentage of PART/TOTAL (rationals accepted)
percent_changeExact percent change FROM_VALUE→TO_VALUE, relative to abs(from_value); refuses from_value=0
calc_statsmean, median, sample stdev, CV (CV > 0.2 = noise swamps the effect)
percentilesp50/p90/p95/p99 by nearest-rank AND interpolation; warns n<100
collision_probabilityBirthday-bound hash collision: 1e5 items/32 bits ≈ 0.69, 1e6/64 ≈ 2.7e-8
data_sizesByte sizes both ways: KiB/MiB (binary) AND KB/MB (decimal)
human_durationHumanised duration + per-day/per-30d rates
epoch_timeEpoch s/ms/µs/ns → ISO 8601 UTC, implausible readings suppressed
radix_convertAny base 2..36, fractions included, non-termination flagged (0.1 base 2)
float_reprWhat binary64 actually stores: exact value, raw bits, ULP, neighbours, representable-or-not
bitsProgrammer-mode integer facts and operations, selected by mode: analysis (was bit_analysis), op (was bitop), widths (was int_widths), repr (was base_repr) — those four former standalone tools were retired in 0.12.0 (CHANGELOG.md)
algebraic_equivAre (a*b)/c and a*(b/c) identical? refactor verification (with float/truncation caveat)
symbolicSymbolic algebra, selected by op: solve (was solve_expression), solve_linear (was solve_linear, name unchanged), simplify (was simplify_expression), limit (was limit_expression) — those four former standalone tools were retired in 0.12.0 (CHANGELOG.md)

Core tools

ToolDescription
list_languages31 languages with extension, compile flag, runtime availability
list_execution_providersExecution-provider identity, interface version, host class, and machine-readable capabilities
execute_codeRun code in any language → stdout/stderr/exit_code/verdict (OK/TLE/MLE/OLE/RTE)/cpu_ms/peak_memory_kb; per-call limits (max_memory_mb, max_output_kb, max_cpu), no_net, compact. With a session and no explicit max_output_kb, oversized output spills into the session workspace (stdout_spill/stderr_spill) instead of just truncating
execute_code_streamProvider-selected execution using the same canonical limits as execute_code, with progress + partial output when the provider supports streaming
trace_executionpython3 only. Runs the same sandboxed executor execute_code uses, plus a per-line event trace (events: line/call/return/exception, changed locals per step) and a static branch/line-coverage report (branches, lines_executed, lines_never_executed) from an AST parse — answers "which lines ran, in what order, and why" rather than just "what did it print"
branch_reachabilitypython3 only. Decides, with z3, which if/elif/else arms and while/for(range, static bounds) loops in ONE function can ever be taken for ANY input — reachable/dead/unknown per branch, a witness when reachable, and boundary_inputs (min/max/equality-edge, via z3 Optimize) shaped for compare_edge_cases's test_inputs. Refuses up front, naming the construct and line, for anything outside + - * // %/and or not/== != < <= > >=/abs min max len on int/bool/str
run_submitSubmit code for background execution; returns a run_id immediately instead of holding the call open
run_inspectPoll a background run: status while running, the full execute_code result shape once terminal
run_cancelCancel a background run; idempotent on an already-terminal run, honest about providers that cannot cancel mid-flight
session_startPersistent session; python3/node get a stateful REPL worker (variables/imports persist across calls), other languages a workspace dir
session_stop / session_listSession lifecycle
session_files / session_read_file / session_write_file / session_delete_fileWorkspace file tools, jailed to the session dir; listings support page_size/cursor, reads return images inline (as_image), and delete removes one file/symlink (never a runner-internal path or a directory) — the in-band recovery when a session is refused for being over CODECALC_MAX_ARTIFACT_COUNT
session_runMulti-file programs: execute an entry file that imports other session files (helper.py, data/...) in the workspace
session_artifactsList files created by executed code (results, images, CSVs)
session_snapshotArchive a session's workspace to a snapshot stored OUTSIDE the jailed workspace (action="save"), or restore one into a new session or, with replace=True, back into the same session (action="restore"); action="list"/"delete" manage them. Files only — never a stateful session's REPL variables. Snapshots die with session_stop unless keep_snapshots=True
install_packageInstall packages (uv pip/npm/gem/go/cargo...) into a session or shared cache
verify_translationProve a port is equivalent: you write the translation, the executor runs both versions on the same inputs and reports match / diverged / inconclusive per input. A pass is graded cross_checked (see Grade vocabulary)
verify_optimizationProve an optimisation: you write the candidate, the executor confirms it still agrees with the original AND times both — accepted only if equivalent and measurably faster. Accepted is graded cross_checked
extract_functionPull a named function + its dependency closure (imports, referenced helpers) into a standalone program and run it (ast-exact for python3, best-effort elsewhere)
compare_edge_casesRun the same logic in N languages on edge-case inputs (empty, zero, negative, float precision) and flag behavioral divergence
convert_unitsDimensional unit conversion via sympy: length, mass, time, speed, energy, power, force, pressure, temperature (°C/°F/K), volume, area, data, frequency
physical_constants22 physical constants with values (c, h, N_A, k_B, G, g, m_e, R, ...)
list_unitsAll 140+ unit aliases for convert_units
evaluate_expressionSymbolic math: integrate(x**2, x), sqrt(144) + 2**10
truth_tableBoolean algebra: a and b or not c, p xor q, a implies b
z3_checkSMT-LIB2 satisfiability + model. An unsat verdict is graded solver_proven; sat is graded ungraded (decided, but not proof-shaped — see Grade vocabulary)
matrixStructured matrix ops: det/inverse/eigenvalues/transpose/rank/trace on a rows array — never a caller string through sympify, so evaluate_expression's [/] RCE screen never applies. Each entry screened individually
analyze_complexityStatic Big-O estimate from code structure, parsed with tree-sitter (every supported language). Reports analysis: tree-sitter|regex-fallback so you can tell a parse from a guess
benchmarkEmpirical Big-O: runs code at increasing N, fits growth curve
compare_executionSame code across N languages side-by-side
runtimes_statusNon-mutating update check: current vs latest for every language runtime, which package manager owns it, and the command that would run
update_runtimesUpdate runtimes. Dry-run by default (apply=False returns the commands); apply=True executes them

Retired tool aliases

The 2026-09-08 merge (0.11.0) folded two lexically-overlapping clusters — flagged by Glama's public review as indistinguishable from a description alone — into one enum-selected tool each, keeping every mode's own parameters, annotations and result shape. The eight old names stayed registered as thin aliases for one minor release so nothing broke mid-upgrade, then were removed in 0.12.0 (CHANGELOG.md). Calling one of them now gets the MCP SDK's own unknown-tool error, not a result:

Retired nameReplacement
bit_analysisbits(mode="analysis")
bitopbits(mode="op")
int_widthsbits(mode="widths")
base_reprbits(mode="repr")
solve_expressionsymbolic(op="solve")
solve_linearsymbolic(op="solve_linear")
simplify_expressionsymbolic(op="simplify")
limit_expressionsymbolic(op="limit")

Grade vocabulary

verify_translation, verify_optimization and z3_check return grade + grade_basis (+ grade_rules_version) on top of their own result. The grade names how strong the evidence for a success actually is; it is derived from evidence those tools already emit, in codecalc/grades.py — the verifiers never assign their own grade.

GradeMeansEmitted by
cross_checkedTwo independently authored programs were both actually run and their outputs agreed. grade_basis names the runtime(s) that did the checking.verify_translation (source vs. port), verify_optimization (original vs. candidate)
solver_provenZ3 returned unsat within its timeout — a machine-checked refutation, not a heuristic. grade_basis names the engine version and the timeout bound. Not sat: see below.z3_check
executedReserved: the claimed computation ran and produced the reported result, with no independent second opinion. Not currently emitted by any tool above — every one of them also clears the cross_checked/solver_proven bar.
ungradedExplicit non-grade for a mismatch, an inconclusive comparison, a rejected optimisation candidate, a measurement failure, a Z3 unknown verdict, and — deliberately — a Z3 sat verdict. A real value on grade, never an absent key. Never a softened stand-in for one of the three grades above.any of the above, on a non-success

z3_check's sat verdicts are graded ungraded, not solver_proven, even though sat is just as decisive a verdict as unsat. The ticket's motivating pattern is proving a property P by asserting not-P and checking unsat; a caller running that pattern who gets sat back has learned P is FALSE, and solver_proven on that result would let a reader who skims grade without result mistake a counterexample for a proof. sat's grade_basis says so explicitly: satisfiability was decided, but solver_proven is reserved for unsat so a counterexample can never wear a proof grade. Widening sat back into solver_proven later is additive; narrowing it after callers depend on the wider behaviour would not be, so this ships narrow now. Full reasoning: codecalc/grades.py's module docstring.

algebraic_equiv is deliberately NOT graded: it compares two expressions via sympy.simplify(a - b) == 0, a CAS transformation rather than a decision procedure with a checkable certificate, and it is one simplifier's opinion rather than two independent implementations agreeing. None of the three grades describes that evidence honestly.

Runtime self-update

Every language is mapped to its package manager, and codecalc can update its own runtimes:

ManagerLanguagesUpdate command
misepython3, node, bun, deno, ruby, go, erlang, elixir, gleam, zig, java, kotlin, sqlite, duckdb, gradlemise up
rustuprust (stable/nightly toolchains)rustup update
swiftlyswiftswiftly update
aptc, c++, fortran, csharp, php, perl, lua, tcl, r, jq, bash, zshapt-get install --only-upgrade (language packages only)
npmtypescript/tscnpm update -g
uvmojouv tool upgrade mojo
nixhaskell (on-demand)nothing persistent

runtimes_status is always safe. update_runtimes refuses to mutate unless apply=True is passed explicitly — and it only touches the package manager that owns each language (never the Rust sandbox, which has no update powers).

One of those managers is elevated: apt updates system packages, so its command starts with sudo. apply=True is an argument a connected model controls, so that branch takes a second key the model does not have — the host must set CODECALC_ALLOW_RUNTIME_APPLY=1. Without it the apt command is reported as skipped with ok: false and the variable named, while the unprivileged managers still run. sudo -n already fails closed where a password is required; this covers the passwordless-sudo rule common on developer machines and CI images, which is exactly where -n does not stop it.

Run the server

cd /path/to/codecalc && .venv/bin/python -m codecalc.server
# stdio transport — register with any MCP client

# The identical tool/resource registry over stateless Streamable HTTP:
.venv/bin/python -m codecalc.server serve-http --host 127.0.0.1 --port 8000

Streamable HTTP binds to loopback by default. Bearer-token auth (CODECALC_HTTP_TOKEN) is required for any non-loopback bind — serve-http refuses to start on a routable address if neither it nor --oauth-issuer (below) is set, and the static-token comparison is constant-time — and optional on loopback, where an MCP client spawning the process is already inside the trust boundary. Setting a token does not change the single-operator threat model: put an authenticating reverse proxy and the stronger process/container isolation described in SECURITY.md in front of it before exposing it beyond one operator's own machine.

For hosted use only, serve-http also accepts --oauth-issuer URL (or CODECALC_OAUTH_ISSUER) as an alternative to the static token — off by default; the static-token path above is unchanged when it is unset, and setting the variable costs nothing outside serve-http itself: doctor, --help, serve-strict, and the bare stdio server never touch the network over it, only serve-http's own startup does. The issuer and JWKS URLs must be https:// unless the host is loopback (for local testing); an issuer that is plain http:// on a real host, or cannot be reached at all, fails serve-http's startup outright with a message on stderr rather than starting a server no token could ever pass. Given a reachable issuer, serve-http validates each bearer token as a JWT against that issuer's own JWKS (RS256/ES256; the JWKS URL is discovered once from <issuer>/.well-known/openid-configuration, or pinned with --oauth-jwks-url) and checks its issuer, audience, expiry, and not-before. It also serves RFC 9728 Protected Resource Metadata at /.well-known/oauth-protected-resource/mcp, and a request with a missing or invalid token gets WWW-Authenticate: Bearer resource_metadata="..." pointing at it, per the MCP authorization spec (2025-06-18 and later). --oauth-audience defaults to this server's own resource URL; --oauth-scopes "s1 s2" requires every named scope on the token, checked by the SDK's own auth middleware. If both a static token and an issuer end up configured at once, the issuer wins — a request bearing the static token's exact value is rejected like any other invalid bearer value, and a warning naming both settings is printed to stderr at startup. codecalc runs no /authorize or /token endpoint of its own (it is a resource server only, never an authorization server), so there is no dynamic client registration surface here either; a 2026-07-28-era client that needs one uses a Client ID Metadata Document against its OWN authorization server, not against codecalc.

Point an MCP client at it:

{ "mcpServers": { "codecalc": { "command": "/path/to/codecalc/.venv/bin/python",
                                "args": ["-m", "codecalc.server"],
                                "env": {
                                  "PYTHONPATH": "/path/to/codecalc",
                                  "CODECALC_RUNTIME_PATH": "/path/to/mise/shims:/usr/local/bin:/usr/bin:/bin"
                                } } } }

MCP protocol

Protocol revision 2026-07-28, on the official mcp SDK 2.0. Not fastmcp: fastmcp 3.x pins mcp>=1.24,<2.0 and so cannot reach this revision at all.

Verifying that is less obvious than it looks. mcp.types.LATEST_PROTOCOL_VERSION reads 2026-07-28 regardless of what a given connection negotiated, and the same server answers on either protocol depending only on how you connect:

clientnegotiatedcache hints
ClientSession.initialize()2025-11-25dropped
Client(..., mode="auto")2026-07-28applied

So tests/test_mcp_protocol.py asserts the negotiated value from a real connection. The legacy path still works — backward compatibility is a feature — it just must not be mistaken for the new protocol.

Worth noting for anyone reading the spec's headline change: 2026-07-28 removes protocol-level sessions, and directs servers needing cross-call state to use "explicit, server-minted handles passed as ordinary tool arguments". That is exactly what codecalc's session_id already is.

The result contract

Every result carries contract_version, currently 1.6.0. The published schema is docs/contract/result-v1.schema.json and the policy behind it — what MAJOR/MINOR/PATCH may change, the twelve-month deprecation window, worked success/failure/timeout examples, and the migration path from unversioned servers — is in docs/contract/README.md.

For in-process Python use, the supported protocol-neutral service boundary—and the session/storage internals that are deliberately not public—is documented in docs/embedding.md.

Two things a caller should know before reading anything else:

  • ok means "ran and exited 0". A program that behaves exactly as intended and exits 3 comes back ok: false, exit_code: 3, verdict: "RTE". To tell a failed program from a failed request, read verdict — a request that never reached a runtime has no verdict at all, and has a code instead.
  • code is the branch target, not error. Eight stable values; the prose in error is free to improve and is not a contract. An unrecognised code must be treated as internal — that is what lets a 1.x client survive a 2.0.0 server, though adding a code is still a MAJOR change, because the published enum is closed and a strict validator rejects the result first.
  • Truncation reports a size, not just a flag. output_truncated says output was cut; stdout_bytes / stderr_bytes say by how much — the bytes the program actually produced, before the cap. A 200 000-character print under max_output_kb=1 returns 1 039 bytes of stdout and stdout_bytes: 200001, so a caller can size a retry instead of guessing. null there means not measured (nothing ran); a program that printed nothing reports 0.

The schema is JSON Schema 2020-12 — the dialect MCP 2026-07-28 defaults tool outputSchema to — so a client can validate our results with it directly. scripts/check_contract.py regenerates it from codecalc/contract.py and fails on a diff, and separately re-derives both backends' verdict vocabularies from main.rs and executor.py: check_parity.py compares the two backends' key sets and is structurally blind to a new verdict value, which would leave the published enum short and make a strictly validating client reject a good result.

Configuration

All optional. codecalc runs with none of these set.

VariableDefaultWhat it does
CODECALC_HTTP_TOKEN(unset)Bearer token for the Streamable HTTP transport (serve-http). Unset, the transport is loopback-only — binding a non-loopback address without this set is refused outright. Set, the token gates every request via a constant-time comparison; stdio ignores this entirely.
CODECALC_HTTP_URLhttp://127.0.0.1:8000What the HTTP transport's auth metadata advertises as its own URL. Only consulted when CODECALC_HTTP_TOKEN is set; the loopback default matches the offline-by-default posture rather than guessing a public one.
CODECALC_OAUTH_ISSUER(unset)Same as --oauth-issuer: validate serve-http bearer tokens as JWTs against this issuer instead of the static CODECALC_HTTP_TOKEN. Off by default. If both end up set, the issuer wins and the static token is rejected — see "Run the server" above.
CODECALC_OAUTH_AUDIENCEthis server's own resource URL (CODECALC_HTTP_URL + /mcp)Same as --oauth-audience: the expected JWT aud claim, and the RFC 8707 resource this server advertises at its own /.well-known/oauth-protected-resource. Only consulted when CODECALC_OAUTH_ISSUER is set.
CODECALC_OAUTH_JWKS_URL(unset) — discovered from <issuer>/.well-known/openid-configurationSame as --oauth-jwks-url: pin the JWKS endpoint instead of discovering it. Only consulted when CODECALC_OAUTH_ISSUER is set.
CODECALC_OAUTH_SCOPES(unset)Same as --oauth-scopes: space-separated scopes a token must carry. Unset, any token that otherwise verifies is accepted regardless of scope. Only consulted when CODECALC_OAUTH_ISSUER is set.
CODECALC_RUNTIME_PATHthe server's own PATH, else /usr/local/bin:/usr/bin:/binThe PATH executed code resolves runtimes on. Set this when an MCP client spawns the server: clients often launch with a stripped environment, so an inherited PATH can miss a toolchain manager's shims entirely and most languages silently become unavailable. list_languages reports what actually resolved.
CODECALC_EXEC_BINbin/codecalc-exec (arch-matched)Override the sandbox binary. Without one, codecalc falls back to a pure-Python executor — list_languages and execute_code still work, but the Rust path is the production one.
CODECALC_REQUIRE_NATIVE(unset)Fail-closed: refuse to start if no usable codecalc-exec binary was found (checked at import, so this is also a server-start check), instead of silently answering every call on the weaker Python fallback. Raises naming CODECALC_REQUIRE_NATIVE and the paths that were checked.
CODECALC_EXECUTION_PROVIDERlocalDefault execution-provider ID. Explicit execute_code(provider=...) selection still wins. Setting this to an unregistered provider fails explicitly; it never falls back.
CODECALC_PISTON_URL(unset)Register the non-local open-source Piston v2 provider at this absolute HTTP(S) base URL. No public service is contacted by default.
CODECALC_PISTON_AUTHORIZATION(unset)Exact value for Piston's Authorization header. It is scoped to the Piston transport and redacted from normalized results, descriptors, health, and receipts.
CODECALC_STRICT_URL(unset)Activate the current OS's <host>-strict provider as an authenticated client of the Linux strict execution service. Without it, strict selection fails closed. The adapter verifies the remote enforcement handshake before sending source.
CODECALC_STRICT_AUTHORIZATION(unset)Exact value for the strict service's Authorization header. It is never published in descriptors, doctor output, errors, or receipts.
CODECALC_RUN_STATE_DIR~/.codecalc/runsDurable metadata-only journal backing run_submit/run_inspect/run_cancel, for every provider (not only managed strict runs). Source, stdin, output, and credentials are never written there. On restart, recorded orphan runs are cancelled and cleaned through their owning provider where it supports that; where it does not (the built-in local provider), there is nothing to signal and the record is simply marked recovered.
CODECALC_MAX_ACTIVE_RUNS64Admission cap for run_submit: how many runs may be running/cancelling at once before further submissions are refused with a resource_exhausted error. Bounds the in-memory run table and its thread pool against an unbounded burst or a caller that never inspects/cancels what it starts. An empty, non-numeric or non-positive value falls back to 64 with a message on stderr — a set-but-empty variable is a shell and compose-file commonplace, and it used to abort the server's import.
CODECALC_ALLOW_RUNTIME_APPLY(unset)Permit update_runtimes(apply=True) to run the elevated update commands (apt, via sudo). Unset, they are skipped with ok: false naming this variable, and the unprivileged managers still run. Deliberately an environment variable rather than a tool argument: apply is something a connected model can flip, and this is not. Accepts 1/true/yes/on; an empty value is not consent.
CODECALC_SESSION_ROOT~/.codecalc/sessionsWhere session workspaces live. Keep this codecalc-private. codecalc cleanup --write --include-unmarked removes plain, session-shaped subdirectories under it on a heuristic (name shape + age) that is a loose filter, not a strong one — never point it at a directory anything else writes into.
CODECALC_CLEANUP_ABANDONED_AGE_HOURS24How old (and untouched) a marker-less, session-shaped directory must be before codecalc cleanup --include-unmarked will consider it abandoned. Only consulted with --include-unmarked; the default cleanup invocation never reads it.
CODECALC_PACKAGE_ALLOWLIST(unset)Deny-by-default allowlist for install_package. Unset, any syntactically valid package name may be installed (today's behaviour). Set, only listed packages install — anything else is refused before any subprocess or network work, with the stable permission_denied code. Comma-separated; each entry is <language>:<name> (scoped to one ecosystem) or a bare <name> (every ecosystem). Matches the bare name, ignoring [extras] and ==version pins.
CODECALC_SESSION_IDLE_TTL_SECONDS(unset)Idle-expiry for stateful (python3/node) session workers: a session untouched for longer than this is reaped — worker killed via the same teardown session_stop uses — on its next access. Unset, a session worker lives until session_stop or server exit, same as before this existed. A subsequent call on an expired session gets ok: false with the stable worker_failure code, never a silent respawn.
CODECALC_SESSION_DISK_QUOTA_MB512Per-session ceiling on total workspace disk. session_write_file and oversized-output spilling refuse BEFORE writing (resource_exhausted, no partial file); code run via execute_code(session_id=...)/session_run is checked before it starts and, since its own writes cannot be pre-checked, again after — an over-quota run still returns its result, now with disk_quota_exceeded plus usage/limit, and the session's next write/run is refused until usage (re-measured fresh each time) drops back under the line. Also the cap a SESSIONLESS run's per-run dependency workdir is held to (codecalc/dependencies.py, checked after each successful install) — reused rather than a second, independently-tunable constant, since it is the same kind of workspace in every way that matters here.
CODECALC_TOTAL_DISK_QUOTA_MB8192Global ceiling on disk summed across every session workspace on this host — closes the gap where staying under the per-session quota by opening many sessions would otherwise be unbounded. Same enforcement points and resource_exhausted contract as CODECALC_SESSION_DISK_QUOTA_MB.
CODECALC_MAX_ARTIFACT_BYTES16777216 (16 MiB)Per-write size ceiling for anything a session write path creates — independent of the total quotas above, so one runaway file cannot hide under a generous session/global total. A WRITE-time cap; distinct from RESOURCE_MAX_BYTES (4 MiB), which caps what a read may serve back.
CODECALC_MAX_ARTIFACT_COUNT500Per-session ceiling on the number of artifact files — catches a session writing one byte at a time into thousands of tiny files, a shape no byte-sized cap alone bounds. Only a write that creates a NEW file is checked; overwriting an existing one always succeeds regardless of the count.
CODECALC_MIN_HOST_FREE_MB256Refuse a session write when the HOST's free disk space drops below this — protects the host even when every quota above is generous, since a shared host can be driven low by something that is not a codecalc session at all. Measured with shutil.disk_usage, which works identically on Windows, unlike statvfs.
CODECALC_MAX_SNAPSHOT_BYTES268435456 (256 MiB)session_snapshot(action="save") refuses to archive a workspace whose files sum to more than this — independent of the SESSION disk quotas above, since a snapshot is written OUTSIDE any session's own workspace and quota.
CODECALC_MAX_SNAPSHOTS_PER_SESSION10Per-session ceiling on the number of snapshots kept at once — catches many small snapshots the byte cap alone would not, the same "count cap alongside the byte cap" shape CODECALC_MAX_ARTIFACT_COUNT already applies to workspace files.
CODECALC_CAPABILITY_POLICY(unset)Capability broker. Unset, no brokering — a job's capabilities run as requested (today's behaviour); the execution receipt still discloses them under provider.capabilities with brokered: false. Set, comma-separated directives narrow them: deny-network forces no_net on a job that did not request network (enforced where the provider can, disclosed as effective where it cannot); allow-network explicitly grants network to a job that requested it; strict rejects a job whose denial the provider cannot enforce. The broker never approves a capability the request did not ask for — an escalation is refused with permission_denied / capability_not_requested, before any side effect.
CODECALC_AUDIT_LOG~/.codecalc/audit/audit.logAppend-only JSON-lines audit stream for broker decisions and security-relevant side effects (denied capability, refused install, cleanup). Each event carries a source-safe timestamp, the run/session id, the decision and reason, and never the executed source or a credential. Set to a path to relocate it; set empty to disable. Best effort — a write failure never fails a run.
CODECALC_PROCESS_HEADROOM512Fork-bomb guard. RLIMIT_NPROC is a uid-wide task budget, not a per-sandbox one — the kernel compares it against every thread your user owns, machine-wide. So codecalc measures the ambient count per execution and sets the limit to ambient + headroom: a bomb can add at most this many tasks, while a runtime wanting a few threads always has room however busy the box is.
CODECALC_MAX_PROCESSES(unset)Escape hatch: pin RLIMIT_NPROC to an absolute value and skip the measurement.

The strict service runs on Linux x86_64 or ARM64 with Docker Engine, cgroup v2, and an explicitly registered gVisor runsc runtime. Its executor image must be pinned by @sha256: digest on the execution path. That image is published to GHCR (ghcr.io/the-40-thieves/codecalc-exec, multi-arch amd64+arm64) by the publish-executor-image workflow, which an operator dispatches (workflow_dispatch); the workflow commits the immutable digest into docker/executor-image.lock, and published_strict_image() resolves it as the production default. Until that first dispatch no digest is pinned and the execution path fails closed — it never falls back to the mutable local diagnostic tag (codecalc-exec:strict), which doctor and the conformance suite keep using. The default systrap platform works without KVM, so the same authenticated service can be used from Linux, macOS, and Windows; strict clients never fall back to native local execution.

Provisioning and running any of the three strict backends in production — the gVisor+Docker host, Windows AppContainer hardening, and the macOS/Windows remote-client configuration — is covered in docs/deployment/README.md, separate from the provider interface itself in docs/contract/provider-v1.md.

Both backends resolve CODECALC_RUNTIME_PATH identically, and scripts/check_parity.py fails CI if the Rust and Python copies of that contract ever drift — including if a machine-specific home directory finds its way back into the default.

Tool-definition token cost

codecalc's tools/list returns 51 definitions. Measured with o200k_base as a proxy on the served JSON, that is 78,586 bytes / 20,630 tokens of descriptions and input schemas (up from 64,643 bytes / 17,277 tokens at the same 51 tools), and every client pays it before the first user message. The number has grown with the descriptions, not the count: the disambiguation sentences and the per-mode text on bits/symbolic are what a selection-accuracy-first server spends its tokens on. The latest jump (+13,943 bytes, +3,353 tokens) is every one of the 152 tool PARAMETERS gaining its own description in the input schema (Annotated[<type>, Field(description=...)]) — the docstrings above did not change, so scripts/tool_select_eval.py's selection-accuracy numbers (it scores only name + docstring, never the input schema) are unaffected by this change.

A follow-up pass then edited the DOCSTRINGS themselves, now that every parameter's own syntax/default/range lives in its schema description and no longer needs restating in prose: measured with the same o200k_base proxy (a mcp.Client.list_tools() dump, by_alias=True, exclude_none=True, one tool per JSON object), 77,237 bytes / 20,439 tokens — down from the 152-parameter figure above (-1,576 bytes, -412 tokens), short of the ~19,000-token target this pass aimed for. What moved which way: the execute_code/execute_code_stream/run_submit/session_run cluster and six other tools (truth_table, session_files, algebraic_equiv, compare_threshold, percentage, percentiles) each gained one disambiguation/usage sentence naming a sibling tool, verify_optimization's docstring was cut to about 60% of its length, and pure schema-restating sentences ("languages = comma-separated subset", explicit default/range call-outs, repeated operation lists) were removed across the file — but several of THOSE removals had to be partly reverted once scripts/tool_select_eval.py showed they deleted discriminating vocabulary BM25 actually leans on (see scripts/data/tool_select_baseline.json, regenerated by this pass: full 151→150/234, dev 128→130/184, core 79→78/116 top-1 hits, every surface within the eval's own DEFAULT_EPSILON_HITS), which is most of why the net reduction is smaller than the additions alone would suggest.

A per-tool icons field (2025-11-25+) was tried and measured, not assumed: one tiny inline data:image/svg+xml;base64,... glyph per tool GROUP, under 300 bytes even for the largest of six — small per icon, but Tool.icons is a per-TOOL field, so each of the tools repeats its group's full base64 payload on the wire, and base64 tokenizes far worse than prose under a BPE encoder. Measured on the full served tools/list payload: +11,540 bytes, +6,665 tokens (o200k_base) — real and non-trivial on a server whose whole pitch (see "Reducing the tool surface" below and docs/design/ 2026-08-10-tool-facade.md) is that tool SELECTION accuracy matters more than saving a few tokens elsewhere. Removed. codecalc's MCPServer still carries one SERVER-level icon plus a website_url — both ride on initialize, once per connection, not once per tool, so they do not touch tools/list at all: measured before/after, the served tools/list payload is byte-identical (59,902 bytes / 15,952 tokens either way, as measured at the time on the then 52-tool surface) — +0 on the number this section exists to track.

codecalc does not hide its tools behind a discovery facade, and that is deliberate: the tool surface is where per-operation approval prompts, audit names and typed schemas live, and collapsing 51 tools into one dispatcher makes install_package and percentage look like the same permission to a client that approves by tool name. The cost is real, but the client is the better place to solve it, because the client can defer definitions without giving up the schemas or the per-tool boundary.

If you are paying too much for codecalc's definitions:

  • Claude Code defers every MCP tool by default — tool search is on by default, with no token floor codecalc needs to clear. auto loads a server's tools upfront only while their definitions total under 10% of the context window and defers all of them once that 10% is reached; false loads everything upfront regardless of size (Claude Code MCP docs, https://code.claude.com/docs/en/mcp, retrieved 2026-09-07). calc_exact, execute_code, verify_translation, verify_optimization, and list_languages carry _meta["anthropic/alwaysLoad"] (per that same doc, "your 3-5 most frequently used tools") so they stay loaded even when a client defers everything else; install_package and update_runtimes carry _meta["anthropic/requiresUserInteraction"], which forces a permission prompt on every call regardless of the session's permission mode — both change the host and both fetch from a registry. execute_code, execute_code_stream, session_run, compare_execution, and run_inspect carry _meta["anthropic/maxResultSizeChars"] = 499520 (2 * 240 KiB + 8_000), the truncation hint for the one tool family whose output can legitimately approach it — run_inspect carries it because its terminal reply, once a run_submit-started run finishes, is the same envelope execute_code returns. This bounds the serialized TEXT content block only — the JSON result as the string a client renders as the tool's reply — not the whole MCP response: every one of these five tools except session_run also declares outputSchema, so the SDK additionally attaches structuredContent with the same JSON ("MCP server developers can configure custom output limits for individual tools by specifying _meta['anthropic/maxResultSizeChars'] in the tool's listing, up to a hard maximum of 500,000 characters", same doc as above, describes the text result specifically) — so a typed tool's total wire payload approaches twice this hint. session_run's inlined artifact content blocks (image/text/link, up to 8 within a 4 MiB encoded budget — see its own docstring) are likewise separate blocks outside this bound. 240 KiB per stream is a hard CEILING max_output_kb is clamped to on every tool that accepts it (execute_code, execute_code_stream, run_submit), separate from the 64 KiB DEFAULT 0 selects — chosen as the largest round-KiB ceiling that keeps the TEXT-block hint under that 500,000-char limit; there is no other ceiling on that parameter today, and raising it further would push the hint over that limit. A run whose real output needs more than 240 KiB belongs in a session instead: leaving max_output_kb at its default with session_id set spills oversized output to a full-fidelity file, readable in full via session_read_file, rather than truncating it. compare_execution takes no max_output_kb of its own, but accepts an unbounded number of snippets, so a many-language comparison can still legitimately exceed the hint.
  • Claude API, via the MCP connector, takes defer_loading once on the toolset's default_config, or per tool in configs. Deferred definitions stay out of the system-prompt prefix, prompt caching is preserved, and a matching tool is expanded into its full definition when the model searches for it.
  • OpenAI's Responses API has the same knob under a different name: defer_loading: true on an MCP server tool definition (OpenAI Responses MCP tool guide, https://developers.openai.com/api/docs/guides/tools-connectors-mcp, retrieved 2026-09-07).
  • VS Code caps a single chat request at 128 enabled tools and groups excess tools behind "virtual tools" above a configurable threshold (VS Code agent tools docs, dated 2026-09-02, https://code.visualstudio.com/docs/copilot/agents/agent-tools). Windsurf / Cascade caps at 100 total tools (Cascade MCP docs, https://docs.devin.ai/desktop/cascade/mcp, retrieved 2026-09-07).
  • The MCP specification itself has no deferral mechanism — no tool search, grouping, tags, or toolsets; a server can only publish ttlMs/cacheScope hints and paginate tools/list (MCP spec 2026-07-28, https://modelcontextprotocol.io/specification/2026-07-28/server/tools, retrieved 2026-09-07). A client without one of the mechanisms above pays the full cost regardless of what codecalc does.
  • Any client can filter which of the 51 tools it exposes to the model. Nothing here requires codecalc to change.

A server-side facade remains under consideration for clients with no such mechanism (docs/design/2026-08-10-tool-facade.md), and is not implemented.

Trimming a description to cut this cost is exactly the change scripts/tool_select_eval.py exists to gate: an offline, labeled eval of whether a deterministic lexical (BM25) selector still picks the right tool for a plain-language ask, scored against the live tools/list text. Measured v1 baseline (196 hand-labeled prompts, none containing their own target tool's name — see the script's own docstring): 60.71% top-1 / 75.51% top-3 accuracy on the full surface (62.75% / 63.0% top-1 on dev / core respectively). It is a lexical proxy, not a model — see the script's module docstring for exactly what a green run does and does not prove.

The checked-in baseline PINS the exact labeled corpus by content hash (prompt_set_sha256); a --baseline compare against a corpus that no longer hashes to it fails with a distinct "corpus changed" error rather than silently scoring a smaller, easier prompt set against the old numbers. And because a tool can be top-1-wrong against full's 51 distractors (zero headroom to lose) while still having real headroom against core's much smaller distractor set, both the regression compare and the ablation self-check (replacing real descriptions with a generic stub, one tool at a time, across every candidate tool — no sampling) run separately against all three of full/dev/core, wired into CI via tests/test_tool_select_eval.py so the gate is proven live, on every surface, on every run — not just at the PR that added it.

BM25 is a lexical proxy, not a model — scripts/tool_select_llm_eval.py is the model-driven half, calling a real chat model over a live gateway with the identical tool catalog and labeled corpus; it is opt-in (workflow_dispatch, advisory rather than a hard gate) rather than wired into every PR, and its measured numbers live in docs/tool-selection-eval.md next to BM25's own.

Reducing the tool surface

For an operator who would rather not configure every client, codecalc also has a first-party knob: CODECALC_TOOLS registers only a chosen slice of the 51-tool surface, so a client that never enables tool search still pays for a smaller tools/list.

On a client with no deferral mechanism of its own, the client's own allow-list does the same job from the other end — OpenAI's allowed_tools, Gemini CLI's includeTools/excludeTools, or Codex CLI's enabled_tools/disabled_tools all narrow what a given session sees without touching the server.

Every tool also now carries a ToolAnnotations hint (readOnlyHint, destructiveHint, idempotentHint, openWorldHint — see codecalc/server.py's GROUP_ANNOTATIONS/TOOL_ANNOTATION_OVERRIDES tables for the value on each of the 51). Codex CLI's writes approval mode (v0.144.0+) reads readOnlyHint directly: a tool marked readOnlyHint: true skips the approval prompt, everything else still asks. That covers the whole calculator group (20/20 pure) plus the read-only members of the mixed groups — list_languages/list_execution_providers/runtimes_status/ branch_reachability in execution, z3_check/algebraic_equiv in verification, session_list/session_files/session_read_file/session_artifacts/ run_inspect in sessions, and analyze_complexity in analysis — without codecalc doing anything client-specific; the annotation is the same hint every MCP client reads, writes just happens to be the mode that consumes it.

This is not the facade the section above declines to build. Every tool a group activates keeps its own name, its own typed input schema and its own per-tool approval prompt — a group that is not active simply never registers its tools with the MCP SDK at all, so they are absent from tools/list and rejected by tools/call, not merely hidden behind a dispatcher a client could still invoke by guessing the name.

Every tool belongs to exactly one group:

GroupTools
calculator (20)calc_exact, compare_threshold, percentage, percent_change, calc_stats, percentiles, collision_probability, data_sizes, human_duration, epoch_time, bits, radix_convert, float_repr, symbolic, convert_units, physical_constants, list_units, evaluate_expression, truth_table, matrix
verification (5)verify_translation, verify_optimization, algebraic_equiv, compare_edge_cases, z3_check
execution (8)list_languages, list_execution_providers, execute_code, execute_code_stream, trace_execution, branch_reachability, compare_execution, runtimes_status
sessions (13)session_start, session_stop, session_list, session_files, session_write_file, session_delete_file, session_read_file, session_run, session_artifacts, session_snapshot, run_submit, run_inspect, run_cancel
analysis (3)analyze_complexity, benchmark, extract_function
admin (2)install_package, update_runtimes

CODECALC_TOOLS takes a comma-separated list of group names, preset names, or both:

PresetExpands to
corecalculator
devcalculator, execution, verification, analysis
fullevery group (the default)
CODECALC_TOOLS=calculator            # just the calculator (20 tools)
CODECALC_TOOLS=core                  # same thing, by preset name
CODECALC_TOOLS=calculator,execution  # two groups, unioned
CODECALC_TOOLS=dev                   # a coding-assistant slice (36 tools)

Unset or empty registers every group — 51 tools, same as today — so nothing changes for an operator who does not set this. An unknown group or preset name is a loud startup failure naming the bad value and every known group/preset, never a silent fallback to "everything" or "nothing": either direction would turn a typo into a footgun nobody notices until it matters. codecalc doctor prints the active groups, the full group→tools mapping, and how many tools this process actually registered, whatever CODECALC_TOOLS is set to.

Client-side deferred loading (the section above) and this env var compose cleanly: point a client with no deferred-loading mechanism at a CODECALC_TOOLS-restricted process, or use both — a smaller declared surface still benefits from being deferred.

Test

Each file is a standalone script that prints one PASS/FAIL line per assertion and exits non-zero if any failed — no test runner, no plugins.

cd /path/to/codecalc

# everything. `|| break` used to be `|| break` alone, which stopped at the
# first failure AND left the loop exiting 0 — a red suite reported success to
# anything wrapping this command. This form runs them all and carries the
# failure out.
fail=0
for f in tests/test_*.py; do PYTHONPATH=. .venv/bin/python "$f" || { echo "FAILED: $f"; fail=1; }; done
for f in scripts/*.py;    do PYTHONPATH=. .venv/bin/python "$f" || { echo "FAILED: $f"; fail=1; }; done
[ "$fail" -eq 0 ]   # the exit status of the whole run

# or individually
PYTHONPATH=. .venv/bin/python tests/test_smoke.py           # every language, via the Rust executor
PYTHONPATH=. .venv/bin/python tests/test_mcp_all.py         # every tool over MCP stdio, answers checked
PYTHONPATH=. .venv/bin/python tests/test_executor_sweep.py  # sandbox regressions

71 test files and 19 CI-invoked scripts, 2184 assertions. "CI-invoked" means referenced by path (scripts/<name>.py) from a job in .github/workflows/*.ymlscripts/check_claims.py derives the count that way and gates it, so a script wired into a workflow without this sentence changing, or this sentence bumped without a workflow change, fails the build. Nothing in the suite needs the internet, so none of it is ever skipped for lack of a network.

It can skip for lack of a capability, and that is correct rather than a regression: a machine without a symlink privilege, without a given language runtime, or without a built native executor cannot exercise the cases that need them. The suite reports three distinct outcomes — the property holds, the property is broken, and this machine cannot exercise it — and every skip names its real cause. A nonzero skip count on Windows or in fallback mode is the healthy result; what would be wrong is a skip reading as a pass.

This paragraph previously claimed zero skips unconditionally. That became false the moment the suite learned to distinguish the third outcome, and nothing gated it: check_claims.py gates the counts below, not the prose around them. The counts are gated by scripts/check_claims.py: they were written by hand once and were stale within three pull requests, which is exactly the failure the rest of that script exists to prevent. Four of the files are regression suites named after the sweep that produced them — test_bug_sweep, test_executor_sweep, test_python_sweep, test_network_modules — and each one's docstring states the defect it locks out and how it was reproduced, because a regression test whose reason has been forgotten is the first one deleted.

Two rules the suite holds itself to, learned from breaking both:

  • Assert the value, not the shape. Three of these files once had no assertions at all: they called tools, printed the output and exited 0. They caught a crash and never a wrong answer — a runtimes_status total replaced with -999 passed, printing total = -999.
  • Don't pin what varies. benchmark and compare_execution rank by measured time, so their winner moves under load; their structure is asserted and their timing is not. runtimes_status is checked against itself — the summary must agree with the data it summarises — so it holds on any machine rather than describing this one.

Platform support

Linux, macOS and Windows. The three do not offer the same primitives, and the executor reports which ones it could not apply in an unenforced array on every result rather than letting a caller assume they all held.

The native table below describes the local provider and is not a hostile-code security boundary. On macOS, <host>-strict instead uses the explicitly configured Linux strict service: the macOS binary performs provider selection, attestation, supervision, and result validation, while untrusted code executes inside the remote cgroup/namespace/seccomp/Landlock boundary. A missing or incomplete service fails before source leaves the Mac and never falls back to native execution.

Symbolic evaluation carries the same idea. Every symbolic tool runs SymPy in a forked child under CPU and memory ceilings with a wall clock the parent enforces, so an expression nobody anticipated is still bounded — SymPy's own maintainers abandoned their attempt at a safe= flag as "security theater", so the screen in safe_expr.py buys time and the child buys the bound. Where there is no fork, the result reports expression_bound_not_enforced_without_fork rather than implying a guarantee.

A second field, output_error, covers the other way a result can be wrong: absent means stdout/stderr are what the program produced, present means at least one of them is not, and names which stream and the OS error. That distinction did not exist until #80 — an output file that could not be read came back as a program that printed nothing, on a run reported as successful. ok now accounts for it on both backends.

GuaranteeLinuxmacOSWindows
Wall-clock timeoutyesyesyes
Kill the whole process treekillpg + PDEATHSIGkillpgTerminateJobObject
Fork-bomb guardRLIMIT_NPROC (uid-wide)RLIMIT_NPROC (uid-wide)Job ActiveProcessLimit, reported unverified
Memory ceilingRLIMIT_ASreported unenforced¹Job ProcessMemoryLimit
CPU-time ceilingRLIMIT_CPURLIMIT_CPUJob PerProcessUserTimeLimit
Open-file ceilingRLIMIT_NOFILERLIMIT_NOFILEreported unenforced
Output capyesyesyes (on read)
no_netseccomp-bpf filter⁶ (falls back to LD_PRELOAD shim²)DYLD_INSERT_LIBRARIES²˒³reported unenforced
Stateful sessionsyesyesyes

¹ Darwin accepts setrlimit(RLIMIT_AS) but does not enforce address space the way Linux does, so setting it would buy an illusion. ² Dynamically-linked programs only — a statically linked binary (Go, by default) ignores it. ⁴ Applied via JOB_OBJECT_LIMIT_PROCESS_TIME, which Windows has supported since XP — this was reported as cpu_limit_unavailable_on_windows until 2026-08-08, and the table said the same, so code and docs agreed with each other and disagreed with Windows. It is not identical to RLIMIT_CPU and the difference is reported rather than glossed: it counts user-mode time only, so a process burning kernel time is not capped by it, and the system checks periodically rather than immediately. Runs on Windows carry cpu_limit_counts_user_time_only_on_windows in unenforced to say so.

³ Weaker still on macOS, in two ways. SIP and the hardened runtime strip DYLD_INSERT_LIBRARIES for protected and hardened-signed binaries (most signed interpreters), and dyld interposing does not reach calls made inside the shared cache where libSystem lives — a program's own connect() is intercepted, a system framework opening a connection internally is not. Treat macOS no_net as a speed bump, never as isolation.

⁶ Linux only. The executor installs a seccomp-bpf filter in the sandboxed child that refuses the socket(AF_INET/AF_INET6) SYSCALL in-kernel — not a libc symbol, so ctypes/dlsym and raw syscall() calls cannot route around it the way they can around the LD_PRELOAD shim. AF_UNIX still works. Falls back to the shim (with its symbol-level bypass, disclosed in unenforced as no_net_best_effort_shim) when the kernel refuses the filter.

Both are exercised by the suite on every platform. The fork-bomb probe measures the EAGAIN boundary precisely but needs os.fork, so it is POSIX-only; a second probe SPAWNS processes instead, which is the portable operation, and pins the ceiling low through CODECALC_MAX_PROCESSES so it costs two dozen short-lived processes rather than walking up to the fallback. Verified to track the limit rather than something incidental: a headroom of 24 bounds it at 22 children and a headroom of 300 bounds it at 298.

⁵ Windows' ActiveProcessLimit is scoped to the job rather than to the uid, so it avoids the failure mode that broke 14 of 31 runtimes on Linux. CodeCalc now supplies that job at process creation, makes it non-nestable with the minimal JOB_OBJECT_UILIMIT_EXITWINDOWS restriction, and allowlists only the three standard I/O handles inherited by the child.

Measured on Windows 11 Pro: 400 of 400 spawns succeeded against a ceiling of 24, reproduced from two unrelated launchers including Task Scheduler. This is not a failed API call — SetInformationJobObject and AssignProcessToJobObject both return success and the correct limit reaches the job. It is topology. ActiveProcessLimit is not one of the limits combined across a nested job chain; those take the most restrictive value, while this one comes from the process's immediate job. A post-creation AssignProcessToJobObject places the child somewhere in that chain rather than at its end: measured, the child's immediate job reported 0x3000 / APL 0 while codecalc's reported 0x230A / APL 24, so codecalc's ceiling was never consulted.

No parent-side Win32 call returns another process's immediate job or its effective ActiveProcessLimit, so this cannot be closed by inspection. Every compatibility run that assigns the child after creation therefore carries process_limit_enforcement_unverified_on_windows in unenforced. Four further strings can each positively prove a failure; none can prove success, so their silence does not imply enforcement.

Creation-time assignment is the default. It was verified on Windows 11 Pro with a direct Python runtime: 23 children succeeded against a total limit of 24 and the next spawn failed with WinError 1816. Runtime launchers that require an inner job now fail rather than silently escaping the limit; configure a direct runtime executable. CODECALC_WIN_JOB_AT_CREATION=0 retains the old path only as an explicitly unverified compatibility escape hatch.

AppContainer security isolation is a DIFFERENT guarantee from the Job Object's resource limits. The Job Object above caps resources — memory, process count, user-mode CPU — and each run names in unenforced which of those did not bind. The optional AppContainer backend adds a security boundary layered on the same creation-time topology: a least-privilege AppContainer profile (CreateAppContainerProfile, no capability SIDs, so no network), launched with SECURITY_CAPABILITIES in the same STARTUPINFOEX attribute list as the job assignment. Access is granted two ways, deliberately split. The sandbox workdir is granted to the run's own AppContainer SID — per-run, so concurrent runs cannot reach each other's workdirs, and it vanishes with the ephemeral directory. The interpreter directory is granted read+execute to the fixed ALL APPLICATION PACKAGES SID (S-1-15-2-1) as an explicit, non-inheritable ACE applied per file across the tree — because a real interpreter's pre-existing files are inheritance-protected and no inheritable grant reaches them. That interpreter grant is persistent and cached (a marker in codecalc's own state dir; the several-thousand-file walk runs once per interpreter): a deliberate trade-off that leaves a read-only ACE, readable by any AppContainer on the machine, on a public interpreter — rather than re-walking every run. The intended property is that a payload cannot read the user profile, write outside its workdir, or reach the network. It is OFF by default (opt in with CODECALC_WIN_APPCONTAINER=1) and fails closed — if profile creation, SID derivation or an ACL grant fails, the launch is refused rather than dropped to an unconfined process. The isolation has been verified on a Windows 11 box (AppContainer SID present, user-profile secrets unreadable, writes confined to the workdir, network denied, ambient privileges reduced to the two benign ones Windows keeps), yet every run that takes this path still emits appcontainer_isolation_unverified_on_windows: a Server-SKU CI runner cannot exhibit AppContainer behaviour, and the guarantee ultimately depends on the deployment's OS and configuration, so the shipped default stays conservatively disclosed rather than claiming a universal proof.

Two things degrade rather than fail on a given platform: languages whose runtime is absent (list_languages reports available: false), and the shell-wrapped plans — gleam and haskell — which need a POSIX shell to scaffold a project and report available: false on Windows outright rather than resolving through a bash that cannot run them. csharp left that set: .NET 10 runs a single .cs file directly, so it is shell-free on every platform.

Reliability tiers

available/status above is a claim about resolution: did this machine find the command on PATH. It is not a claim about reliability: has codecalc's own CI ever actually run this language and checked the output. The two are orthogonal, and they disagree in practice — a review's own smoke test found the rust and csharp host toolchains failing on a machine where both rustc and dotnet resolved cleanly. list_languages, runtimes_status, and codecalc doctor all report a tier alongside resolution to make that gap visible instead of silent:

TierMeaning
testedA CI job genuinely executes this language and asserts on its real output, on every PR. Currently python3, node, rust, and go — kept deliberately conservative, and gated by scripts/check_runtime_tiers.py so a language cannot claim it without a CI check backing it, or silently drop out of CI while still claiming it. python3/node earn it from the stateful-worker sweep (all three OS legs); rust/go from tests/test_tier_evidence.py, which compiles and runs a real program in each and asserts a per-run computed stdout — on the Linux leg, where skips are promoted to failures. The tier claims "CI executes this on every PR", not per-platform coverage.
best_effortDeclared, with a local smoke fixture (tests/test_smoke.py), and plausibly works on a normal install with the right toolchain — but no CI job runs it, so nothing would notice it silently breaking. Every other language, including csharp, java, and the rest.
plan_onlyA registry entry never validated on any runner, anywhere, not even locally. None today.

codecalc doctor's text output prints both axes side by side rather than folding tier into the resolution summary, so a best_effort runtime that happens to be installed on your machine reads as exactly what it is: resolved, unverified by codecalc, may be broken.

Sandbox guarantees

  • Fresh temp dir per run, deleted on exit (source + binaries + outputs). The deletion is identity-checked: the directory's device and inode are recorded at creation and re-checked before removal, because executed code runs with that directory as its cwd and can rename another one into its place. A caller-supplied --workdir is a session workspace and is never deleted. If the filesystem supplies no file index to identify the directory by, the deletion is refused rather than performed unverified, so temp directories accumulate there instead of the wrong one being removed. That trade is stated because it is the one this guarantee actually makes: it was previously implemented in the Rust executor only, and the Python fallback deleted unconditionally, which CI caught on Windows.
  • rlimits: CPU (timeout+8s), address space 2TiB (V8/JVM need huge VA), file size 256MiB, 256 FDs, core dumps off
  • The timeout is a total budget: compile and run share it, so --timeout 10 cannot take twenty seconds. duration_ms is the run alone; compile_ms and total_ms are reported separately.
  • Wall-clock timeout kills the whole process group (SIGKILL). So does SIGTERM to the executor — PR_SET_PDEATHSIG reaches only the direct child, so a group kill is what covers its descendants, and the executor is the only participant that knows the group id.
  • Output capped at 64KiB per stream, on every path including stateful sessions. Exceeding it is reported as OLE, and the file-size rlimit is kept strictly above the cap so that overflow stays detectable — tying the two together turned a truncated 4MB output into a silent verdict: OK.
  • Fork-bomb guard via RLIMIT_NPROC, sized from the measured ambient task count plus headroom rather than a fixed number. This is a mitigation, not isolation: the budget is shared with every other process your user owns, so concurrent executions draw on the same pool. cgroup v2 pids.max is the real per-sandbox answer and needs delegated cgroup access a stdio MCP server cannot assume — reach for it when this moves behind a container.
  • no_net blocks the network, not every socket: it refuses AF_INET and AF_INET6 and forwards everything else, so AF_UNIX local IPC keeps working.
  • No network namespace isolation (single-host tool; containerize for untrusted code). Note, 2026-09-07: this bullet describes the pre-#242 state. Since #242 (2026-08-21), Linux additionally enforces no_net in-kernel via a seccomp-bpf filter — see the no_net row in SECURITY.md's "Known limitations" table for the current per-platform breakdown. A full network namespace is still only the strict (gVisor) backend's job; this note does not change that.
  • Every result carries a backend field ("rust" or "python") so a caller never has to infer which sandbox actually ran from an absent key — that was possible to confuse with an older build that never reported it at all. The pure-Python fallback cannot provide everything above: it has no no_net shim (reported in unenforced, not silently dropped), and peak_memory_kb comes back None rather than a number, because ru_maxrss is a process-lifetime high-water mark this path has no way to attribute to one run. CODECALC_REQUIRE_NATIVE=1 turns "running on the fallback" into a startup failure instead of a guarantee you have to notice was quietly weaker.

Sessions

A session is a persistent workspace; python3 and node additionally get a long-lived REPL worker so variables and imports survive between calls. What that does and does not buy you:

workspace sessionstateful worker
Fresh sandboxed process per callyesno — one worker serves every call
max_memory_mb / max_cpu / no_netappliedreported in unenforced
RLIMIT_AS / NPROC / FSIZE / NOFILEper callapplied once, at worker start
Output cap + OLEyesyes
Per-call wall clockyesyes — a worker that blows it is killed

A worker cannot take a per-call rlimit after the fact, and --no-net is decided at exec time. Rather than accept those arguments and drop them, the result lists them in unenforced — the same field the executor already uses to say "asked for, not applied". Omit session_id, or use a workspace session, when a ceiling has to be real.

The worker protocol does not share a file descriptor with executed code, and every response carries the id of the request it answers. Both matter: sys.stdout is a Python-level rebind that a subprocess writes straight past, and a corrupted stream that is not resynchronised returns every later call the previous call's result — a well-formed answer to a different question.

The channel differs by platform and the guarantee does not. POSIX hands the worker an out-of-band pipe; Windows has neither pass_fds nor preexec_fn, so the worker appends responses to a file whose path arrives in the environment. Either way a child spawned with inherited stdio writes to fd 1 and cannot reach the protocol. Tests force the file-backed channel on every platform, because an unexercised fallback is one that works until it is needed.

Local operations: status & cleanup

Two CLI-only commands for an operator running a long-lived server, not MCP tools — they don't count toward the tool surface above:

codecalc status          # read-only snapshot: sessions, disk usage, quotas, audit log
codecalc status --json   # the same report, for scripts

codecalc cleanup                          # DRY RUN (the default) — marker-based only
codecalc cleanup --write                  # actually removes marker-based candidates
codecalc cleanup --write --include-unmarked   # ALSO sweep old, unmarked, session-shaped dirs

status reports SESSION_ROOT, how many sessions exist and which of them are idle-expired (the on-disk .codecalc-session-expired marker the idle-TTL reaping leaves behind), per-session and global workspace disk usage, the configured disk quotas and current headroom, the audit log's path and size, and a one-line runtime reliability-tier summary. It changes nothing — no session is started, stopped, reaped, or written to.

cleanup reclaims disk from session directories under SESSION_ROOT. --dry-run is the default — nothing is removed until --write is passed. Because cleanup runs as a SEPARATE process from any server that may be using SESSION_ROOT right now, it has none of that server's in-memory bookkeeping to consult — only what is on disk.

Directory mtime is deliberately NOT trusted as a liveness signal. An earlier version of this feature did trust it, and an adversarial review proved that wrong live: a REPL worker doing purely in-memory work touches no file at all, and even an in-place file overwrite bumps only that file's own mtime, never its parent directory's — so a genuinely-active worker session can look, by directory mtime alone, identical to an abandoned one. The real signal is a per-worker-session liveness lockfile: the server writes its own pid into the session directory the moment a stateful (python3/node) worker starts, and removes it the moment that worker is actually gone (reaped or session_stop). cleanup checks this for every candidate and refuses outright — regardless of marker, age, or the mtime floor below — whenever the lockfile names a pid that is still alive. That is what makes a session any running codecalc server is using is never deleted true for worker sessions.

By default, cleanup considers ONLY directories carrying the idle-expiry marker — the risk-free path, since a marker only ever exists after sessions.py's own idle-TTL reaper has already closed that specific worker for good (session ids are never reused). --include-unmarked additionally sweeps old (CODECALC_CLEANUP_ABANDONED_AGE_HOURS, default 24h), session-shaped, marker-less directories — the one path with real residual risk, because a workspace-only session (no worker) never gets a lockfile to check against, so this path falls back to age + a hard recency floor (nothing modified in the last few minutes is ever touched) as a heuristic, not a proof. Turn it on deliberately, and never point CODECALC_SESSION_ROOT at anything but a codecalc-private directory — the "looks like a session dir" name filter is loose, not strict.

Other safety properties, unconditional on every path: only a direct child of SESSION_ROOT is ever a candidate (never SESSION_ROOT itself); a symlink there is refused, never followed; and removal itself is identity-checked (device/inode, re-verified immediately before the delete) the same way session_stop's own workspace teardown is, so a directory swapped out from under a stale scan is refused rather than deleted.

Language list

python3, node, bun, deno, typescript, ruby, php, perl, lua, tcl, r, elixir, erlang, bash, zsh, mojo, swift, c, cpp/c++, rust, go, fortran, zig, java, kotlin, csharp, gleam, haskell, sqlite, jq, awk — 31 runtimes.

codecalc does not install any of them. It runs whatever is already on CODECALC_RUNTIME_PATH, and list_languages probes each one and reports which actually resolved, so a minimal machine degrades to the subset it has rather than failing opaquely.

Notes

  • Java uses single-file source launch (JEP 330). Kotlin compiles to a jar.
  • gleam/haskell scaffold a temp project (gleam new / nix-shell); csharp runs the file directly (.NET 10 file-based apps).
  • benchmark uses the stdin-N contract: code reads N from stdin, work sized by N.

CI

Five workflows, each documented inline with what it gates and — where a tool was considered and rejected — why it is not there.

WorkflowGates
ci-rustclippy -D warnings; the executor's JSON contract, asserted by running the built binary (OK/TLE/OLE/unknown-language) and confirming a canary secret in the executor's own env does not reach executed code; both static musl cross-builds, checked with file for static linkage; blocknet.so built -Werror, symbol-checked, and confirmed to actually block an outbound connection
ci-pythonruff at a genuine zero residual (ruleset and every exception in pyproject.toml, each with a reason); calc parity on 3.11 and 3.14; the security suite against the Rust backend, with an assertion that the Rust backend is the one under test; MCP stdio round-trip
ci-securityscripts/check_no_eval.py (the CRITICAL-01 invariant), scripts/check_parity.py (the three security constants duplicated in Rust and Python must match), scripts/check_claims.py (README counts and licence), actionlint, gitleaks, trufflehog, osv-scanner, cargo-deny, cargo-audit, and opengrep on a schedule
ci-qualitytypos. Not shellcheck — the repo's last shell script was removed with executor/zig-cc.sh, so the gate would have matched zero files and reported success for scanning nothing; actionlint in ci-security shellchecks every embedded run: block instead. The workflow says so inline.
dcoSigned-off-by on every non-merge commit

Two conventions run through all of them, both borrowed from harder-won experience:

  • Actions are pinned by commit SHA and downloaded tools by SHA-256. A tag is mutable; a digest is not.
  • Every scan asserts it scanned something. A linter pointed at a renamed directory, a dependency scanner with no lockfile to read, and a clean repo all produce the same output — exit 0. Each gate counts its inputs first and fails if the count is implausible.

Licence

Apache-2.0. See LICENSE.

Contributions require a DCO sign-off (git commit -s); dco.yml enforces it.

Keywords

agent

FAQs

Related posts