New:Microsoft Teams Notifications Are Now Available in Socket.Learn more →
Get Started

astra-sweetspot

Package Overview
Dependencies
Maintainers
1
Versions
3
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

astra-sweetspot

Astra vs Sol on real bugs: inspect the receipts, then reproduce one run with Codex.

latest
Source
npmnpm
Version
0.1.2
Version published
Maintainers
1
Created
Source

Astra Sweetspot

Which Astra reasoning effort is worth it? Start with the patches, not a recommendation.

Small, reproducible Astra vs Sol experiments on real public bugs. Independent regression checks, wall time, tokens, and every candidate patch. Built with Codex after reading users' effort questions and Sol vs Astra questions.

See the results → · Method · Suggest a real task

Wall times for eight runs on two historical p-retry bugs. One attempt per model and effort; all passed their focused checks.

npx astra-sweetspot

That command shows bundled results. It makes no model call and needs no API key. To install the command permanently: npm install -g astra-sweetspot.

Launch-day pilot

All 8 runs passed their focused regression checks. On these two tasks, higher Astra effort took more time with the same check score. This does not establish comparative reliability or a universal best effort.

TaskModel / effortChecksSecondsInputCached subsetOutput
abort-delaySol medium7/7130.8124,035106,3684,774
abort-delayAstra low7/778.273,82064,2561,822
abort-delayAstra medium7/7101.4103,40193,0562,284
abort-delayAstra high7/7163.0127,331113,2803,865
abort-cleanupSol medium8/8129.4104,64369,1205,425
abort-cleanupAstra low8/876.283,83174,6241,501
abort-cleanupAstra medium8/877.383,12662,2081,536
abort-cleanupAstra high8/8123.2102,32291,5202,920

Two historical p-retry bugs, four model/effort settings, one attempt per condition. Same starting source and prompt within each task. Both are small tasks from one library and may be in training data. This is not a held-out benchmark or a model leaderboard.

Supplementary upstream runtime checks: all four settings match their starting code's results—63/63 on one task, 12/13 on the other. The one failure also occurs with the known upstream fix; its log and explanation are included. No model trial was rerun.

Input includes cached input. Tokens are not subscription quota or money paid. Time includes Codex startup and its own tests. The table uses requested model identifiers; the CLI may not report the served identifier. Failed and timed-out trials stay in the data.

v0.1.1 telemetry correction: the JSON receipts now also retain the CLI's reported reasoning-token subset and cache-write tokens. These fields were recovered from the eight original usage events; no trial was rerun. See the correction and evidence.

v0.1.2 aggregation fix: if any completed turn omits a usage counter or reports an invalid value, that run's counter stays null instead of showing a partial sum. Explicit zero remains zero. The eight bundled runs each have one complete usage event, so their results are unchanged. Details and regression cases.

Reproduce one trial

npx astra-sweetspot run abort-delay --model astra --effort medium

Requires Node 20+, Git, Codex CLI 0.153.0+, an existing Codex login, and access to the requested model. run invokes a model and consumes your Codex usage. It copies the public fixture into a fresh local workspace and saves a receipt, patch, and private logs under ~/.astra-sweetspot/runs/. Nothing is uploaded automatically. Your current project is not used or changed.

Cases: abort-delay, abort-cleanup. Models: astra, sol. Efforts: low, medium, high. Default timeout: 180 seconds; use --timeout 300 to change it. The published pilot used 180 seconds throughout.

# Inspect the bundled data as JSON
npx astra-sweetspot report --json

# Grade any candidate against the same pristine external checks, without inference
npx astra-sweetspot grade abort-delay /path/to/index.js

What makes a result inspectable?

  • The grader fails the original bug and passes the known upstream fix. npm test verifies this.
  • Only the candidate's implementation is copied into a pristine fixture. Editing local tests does not improve its grade.
  • Every receipt includes source, prompt, grader, candidate, and patch hashes, plus environment and execution status.
  • The prompts, vendored fixtures, grader, candidate patches, and limitations are public. Raw private logs are not uploaded.

Read the method and limitations. Passing these focused checks is not equivalent to passing the entire upstream test suite or proving a production-ready patch.

What should we try next?

Suggest a real task where your current model or effort struggles, with public starting code and a way to check success. Cross-language work, an unfamiliar repository, or a failing integration test would broaden this small sample. Open a task suggestion.

If these receipts help you decide what to try, star the repository to follow the next experiment. Stars indicate interest; they do not prove the tool is being used.

Development: no runtime dependencies or install scripts. Run npm test and npm run site. Model inference is intentionally excluded from CI. Automated checks cover Node 20/22 on Linux and Node 22 on Windows; the published live model trials ran on macOS only.

The shareable chart is also available as SVG. To regenerate it from results/pilot.json, run python3 scripts/plot-pilot.py with Matplotlib 3.10+. This optional plotting dependency is not needed by the CLI or site build.

MIT. Third-party licenses. Independent of OpenAI and p-retry's maintainers.

Keywords

astra

FAQs

Package last updated on 06 Sep 2026

Related posts