
Research
/Security News
Popular npm Packages in the keyv and Cacheable Namespaces Compromised in Active Supply Chain Attack
Popular npm packages keyv and cacheable compromised.
TypeScript evals for LLM apps. Define evals with a typed API, run them through vitest, store every run in SQLite, browse them in a UI, diff them from the CLI.
defineEval — typed API for datasets, tasks, named scorersrun, dev, ui, show, diff, list — all read-side commands emit JSON--note "switched to haiku-4.5") and deltas are computed against the previous tagged versionnpm install --save-dev tsevals vitest
[!NOTE] Requires Node 22+ (uses the built-in
node:sqlitemodule).vitestis a peer dependency.
Create an eval file ending in .eval.ts and export default an eval definition:
// examples/sentiment.eval.ts
import { defineEval } from "tsevals";
export default defineEval<string, "positive" | "negative" | "neutral">({
name: "sentiment",
data: () => [
{ input: "I love this!", expected: "positive" },
{ input: "Worst purchase ever.", expected: "negative" },
{ input: "It's fine, I guess.", expected: "neutral" },
],
task: async (input) => {
// your model / agent / pipeline
return await classifySentiment(input);
},
scorers: {
exactMatch: ({ output, expected }) => (output === expected ? 1 : 0),
llmJudge: async ({ output, expected }) => ({
score: await judge(output, expected),
metadata: { rationale: "..." },
}),
},
});
Run them:
npx tsevals run
Open the UI:
npx tsevals dev # watcher + UI on http://localhost:3939
defineEval(config)defineEval<TInput, TOutput>({
name: string,
data: () => DataItem<TInput, TOutput>[] | Promise<...>,
task: (input: TInput) => TOutput | Promise<TOutput>,
scorers: Record<string, Scorer<TInput, TOutput>>,
})
scorers is a record, so each scorer has a stable identity across runs (used for per-scorer deltas).number or { score: number, metadata?: unknown }. Metadata is stored per row and shown in the UI on click.trialCount: optional integer. Re-runs the full task+scorers pipeline N times per row and averages the score. Use when the task or scorers carry sampling noise (LLM-as-judge, temperature > 0). Per-trial values are stored alongside the mean and surfaced in the UI.**/*.eval.{ts,tsx,mts,...} are picked up by tsevals run.default.*.test.ts — vitest's normal test runner ignores .eval.ts files.tsevals run [pattern] [--watch] [--note "..."] [--json]
tsevals dev [--port]
tsevals ui [--port]
tsevals show <id|latest|prev-version> [--full]
tsevals diff <from> [to=latest]
tsevals list [--limit N] [--versions]
| Command | Description |
|---|---|
run | Run all evals (or a name regex). Saves a row to SQLite. |
run --watch | Vitest watch mode — re-runs on file change. |
run --note "..." | Tag this run as a version with a description. |
run --json | Emit a structured run summary to stdout (no TTY noise). |
dev | UI server + file watcher + auto-rerun + live UI polling. |
ui | UI server only (production / inspect-only). |
show <ref> | Print a run as JSON. --full includes per-row data. |
diff <from> [to] | Per-eval and per-scorer score deltas. Exits 1 on regression. |
list | Recent runs as JSON. --versions for tagged-only. |
Refs latest and prev-version work everywhere a runId is accepted.
Every read-style command emits JSON, exit codes are meaningful, and the loop is scriptable:
tsevals run --json | jq '.score' # post-change score
tsevals diff prev-version || revert_changes # auto-revert on regression
tsevals show latest --full | jq '.evals[].results[]' # inspect rows
A skill for AI coding agents ships at skills/tsevals/SKILL.md. Point your agent (Claude Code, Cursor, etc.) at it for the iteration workflow — when to tag versions, how to inspect regressions, useful jq snippets.
CLI exit codes:
run — 0 if all rows passed, 1 if any faileddiff — 0 if no scorer regressed (delta > -0.001), 1 otherwiseshow — 0 on success, 2 if the ref is not foundUse diff against a named version to fail the build on regression:
# after a green run on main:
tsevals run --note "release-2.4"
# in PR CI:
tsevals diff release-2.4
# exit 0 = no scorer regressed
# exit 1 = at least one scorer dropped
prev-version works the same way against whatever the latest tagged run happens to be.
[!NOTE] LLM-based scorers carry sampling noise. Use
trialCounton the eval definition (see API) to average across multiple trials before relying on a single delta.
Optional. Drop a tsevals.config.{ts,mts,mjs,js,json} in your project root.
// tsevals.config.ts
import { defineConfig } from "tsevals";
export default defineConfig({
dbPath: ".tsevals/runs.db",
});
Currently supported keys:
| Key | Default | Notes |
|---|---|---|
dbPath | .tsevals/runs.db | Where the SQLite history is stored. Relative paths resolve from the config file's directory. |
.ts configs are loaded via jiti so you can use TypeScript syntax without a build step. .mjs / .js use native ESM import; .json is parsed directly.
Each run produces a row in .tsevals/runs.db (SQLite, schema-migrated automatically):
runs row: id, started/finished timestamps, duration, optional noteeval_results row per (data row × eval), with input/output/expected/scores/durationA run with a non-empty note is a version. The UI's score chart and diff prev-version use versions as the comparison baseline.
Tag at runtime with --note "...", or after the fact via the inline note editor on each run in the UI.
.tsevals/runs.db in the working directory (gitignored by default)node:sqlite (Node 22+ built-in, zero native deps)sqlite3 .tsevals/runs.dbThe defineEval shape and the vitest-reporter approach are heavily inspired by evalite. Go check it out.
FAQs
TypeScript evals: typed API, vitest reporter, SQLite history, UI.
We found that tsevals demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.
Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Research
/Security News
Popular npm packages keyv and cacheable compromised.

Security News
A misconfiguration gave three Anthropic models internet access, and one, believing it was in a simulation, shipped a credential-stealing package to PyPI.

Security News
/Company News
Socket has joined the new Composer and Packagist sponsorship program as a launch sponsor, supporting the team that keeps PHP's package ecosystem secure.