
Company News
Socket Joins New OpenJS Program to Fund Node.js Security Work
Socket is joining the OpenJS Security Stewardship Program to fund Node.js vulnerability research, maintainer remediation, and security releases.
External eval harness that proves your product does what it claims — deterministic L0 checks + LLM-judged L2 rubric scoring
The eval harness that proves your product does what it claims.
Val drives a target application (CLI, HTTP), captures a transcript of real behavior, checks it deterministically (L0), and has an independent LLM judge grade it against the product's claims via a rubric (L2). Output doubles as evidence: a hash-chained, replayable record of "we claimed X; here is a machine-verified run demonstrating X."
See the project roadmap for authoritative implementation status and planned work.
npm install -g @stlw/val
Scenario (YAML) Val harness Judge (LLM)
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ claim: "..." │──────▶│ L0 checks │──────▶│ rubric │
│ steps: [...] │ │ transcript │ │ verdict │
│ rubric: [...] │ │ captures │ │ overall │
└──────────────┘ └──────────────┘ └──────────────┘
│
┌──────▼──────┐
│ report │
│ ledger │
└─────────────┘
L0 — Deterministic assertions on every step (exit codes, JSON shape, regex, negation). Free, runs on every commit.
L2 — An independent LLM judge scores the transcript against the product's claims via a rubric. Provider-agnostic — works with OpenAI, Anthropic, DeepSeek, Ollama, or any OpenAI-compatible endpoint.
my-product.yaml):id: my-product
product: my-app
claim: "The CLI exits cleanly with a helpful message on bad input."
steps:
- name: store command
run: my-cli store --content "hello"
expect:
- { exitCode: 0 }
- { jsonPath: "$.id", exists: true }
capture: { id1: "json:$.id" }
- name: retrieve by id
run: "my-cli retrieve ${captures.id1}"
expect:
- { exitCode: 0 }
- { stdoutContains: "hello" }
- name: bad input should fail
run: my-cli store
expect:
- { not: { exitCode: 0 } }
- { stderrContains: "--content" }
rubric:
- id: clean-errors
criterion: "Error messages are human-readable — no stack traces."
weight: high
- id: captures-work
criterion: "Values captured from one step are correctly used in later steps."
weight: medium
judge: { threshold: 0.7 }
val run my-product.yaml --judge remote
cat val-report/report.md
| Mode | Flag | Description |
|---|---|---|
| Remote | --judge remote | Live LLM via any OpenAI-compatible endpoint |
| TypeSafe | --judge typesafe | Jev score judgments with confidence and probability evidence |
| Mock | --judge mock | Deterministic mock — always passes (use VAL_MOCK_JUDGE_FAIL=id to force failures) |
| None | --judge none | Skip L2 entirely — L0 checks only |
# OpenAI
val run scenarios/ --judge remote --judge-base-url https://api.openai.com/v1 --judge-api-key $OPENAI_API_KEY --judge-model gpt-4o
# Anthropic (Claude)
val run scenarios/ --judge remote --judge-base-url https://api.anthropic.com/v1 --judge-api-key $ANTHROPIC_API_KEY --judge-model claude-sonnet-4-20250514
# Local model via Ollama
val run scenarios/ --judge remote --judge-base-url http://localhost:11434/v1 --judge-model llama3
# Deterministic CI
val run scenarios/ --judge mock --ledger val-ledger.jsonl
# TypeSafe Jev (requires TYPESAFE_API_KEY)
val run demo/jev/claim.yaml --judge typesafe --judge-model jev-latest
This README is a concise capability index. The scenario reference is the source for exact YAML semantics, defaults, template merge rules, scenario discovery, and complete examples.
| Top-level field | Purpose |
|---|---|
id | Optional stable scenario identifier (otherwise derived from the file path) |
product, claim | Required product name and behavior claim |
tags | Labels used by --filter |
env, files | Environment variables and files created in the isolated scenario workdir |
setup, teardown | Shell commands before and after the scenario |
timeoutMs | Total scenario time budget |
mode | steps (default) or agent |
goal, docs, maxSteps | Agent-mode goal, product guidance, and step limit |
steps | Shell or HTTP interactions for step mode |
rubric, judge | L2 criteria and per-scenario judge configuration |
auth | Headers injected into later HTTP steps |
extends, imports | Reusable scenario-template references |
| Step field | Purpose |
|---|---|
name, description | Optional human-readable labels |
run | A shell command (exactly one of run or http) |
http | An HTTP request: method, url, optional headers and body |
timeoutMs | Per-step time limit |
expect | L0 assertions |
capture | Named values extracted from the step result |
allowFailure | Record a failed step and continue the scenario |
Use ${name} interpolation in supported scenario strings. The built-ins are ${workdir} (the isolated scenario directory), ${cwd} (the directory from which Val was invoked), and ${target} (the temporary target checkout when --target-ref is used). A capture named id is available to later values as ${captures.id}.
capture maps names to one of these expressions:
| Expression | Captured value |
|---|---|
json:<path> | Value at a JSON path in stdout, for example json:$.id |
regex:<pattern> | First capture group, or whole match, from stdout or stderr |
stdout: | Trimmed stdout |
Val supports two step adapters:
Shell — runs bash -c <command> against the target CLI:
steps:
- run: my-cli store --actor alice --content "hello"
expect: [{ exitCode: 0 }, { jsonPath: "$.id", exists: true }]
capture: { id: "json:$.id" }
HTTP — drives REST APIs via fetch:
steps:
- http:
method: POST
url: "http://localhost:3000/items"
body: { name: "widget" }
expect: [{ status: 201 }, { jsonPath: "$.id", exists: true }]
capture: { itemId: "json:$.id" }
L0 checks available on every step:
| Expectation | Applies to | Description |
|---|---|---|
{ exitCode: n } | shell | Exit code equals n |
{ status: n } | http | HTTP status equals n |
{ stdoutContains: s } | both | Stdout contains substring |
{ stderrContains: s } | both | Stderr contains substring |
{ outputContains: s } | both | Combined output contains substring |
{ outputMatches: re } | both | Combined output matches regex |
{ jsonPath: p, ... } | both | Stdout parsed as JSON; equals, exists: true, gte, lte |
{ maxDurationMs: n } | both | Step completed in ≤ n ms |
{ fileExists: path } | both | File exists in the scenario workdir (path supports ${...} interpolation) |
{ fileContains: { path, contains } } | both | File in workdir contains substring |
{ not: <expectation> } | both | Negates any expectation (nestable) |
A timed-out step fails all its expectations. File paths resolve against the scenario workdir; paths escaping it fail the check.
Cap an entire scenario (setup + steps) with timeoutMs:
id: slow-suite
product: my-app
claim: "The full flow completes quickly."
timeoutMs: 30000
steps:
- run: my-cli full-flow
Each command's timeout is capped to the remaining budget. On expiry the scenario fails with scenario timed out after Nms; teardown still runs.
Reuse config across scenarios with extends (single-inheritance) and imports (compose multiple blocks):
# profiles/base.yaml
product: my-app
tags: [smoke]
env:
DB_URL: sqlite://test.db
setup:
- my-cli provision --db-url "${DB_URL}"
# scenarios/login.yaml
id: login
claim: "Users can authenticate."
extends: ../profiles/base.yaml
imports:
auth: ../profiles/auth.yaml
steps:
- run: my-cli login --user alice
Merge rules: scalars override, objects deep-merge, arrays append (setup/teardown/tags). Steps come from the scenario file — templates provide env, files, setup, teardown, tags, auth, and judge config. Child rubric replaces parent rubric entirely. Template files live anywhere; they're loaded by relative path from the scenario file. Cycle detection prevents circular references.
Declare an auth block to auto-inject headers into HTTP steps after a capture is available:
auth:
inject:
headers:
Authorization: "Bearer ${captures.token}"
After any step captures captures.token, subsequent HTTP steps automatically get the Authorization header. Step-level headers override injected headers. Auth is inactive before the capture is established (no error).
Let Val drive the product like a user, generating steps from a goal using the same LLM-backed judge endpoint:
id: explore
product: my-app
claim: "Users can create and retrieve items."
mode: agent
goal: "Create an item with content 'hello' and retrieve it by the returned ID."
docs: |
Create: POST /items with { content } → { id, content }
Get: GET /items/:id → { id, content }
maxSteps: 8
rubric:
- id: crud-works
criterion: "An item was created and successfully retrieved."
weight: high
The agent receives shell + HTTP tools, observes outputs, and decides next steps. maxSteps (default 20) caps iterations. timeoutMs is checked only between agent iterations, so a single LLM call or action can exceed it; expiry does not currently mark the scenario timed out. After a successful run, --freeze outputs the executed steps as YAML for deterministic replay.
Run scenarios concurrently with --parallel <n>:
val run scenarios/ --parallel 4 --bail
Results are reported in definition order. --bail stops launching new scenarios on first failure and waits for in-flight ones to drain.
Evaluate scenarios against another branch, tag, or commit of the current Git repository without switching the caller's checkout:
val run scenarios/ --target-ref main --judge mock
Val resolves the ref locally first. If it is missing and an origin remote exists, Val fetches only that requested ref. It then creates one temporary detached Git worktree for the run and exposes its path as ${target} and VAL_TARGET_DIR.
Scenarios remain responsible for preparing the target. Val does not guess the package manager or build command:
setup:
- npm --prefix "${target}" ci
- npm --prefix "${target}" run build
steps:
- run: node "${target}/dist/cli.js" --version
Scenario commands still run in their isolated scenario workdirs. The target checkout is removed after the run, including when scenarios fail. --keep-workdir applies only to scenario workdirs.
In CI, the default shallow actions/checkout may not contain the requested ref. Val's narrow origin fallback handles branches and tags available on the remote; configure checkout history explicitly when testing arbitrary commit IDs.
Every run can write a hash-chained, append-only record to val-ledger.jsonl:
val run scenarios/ --ledger
val verify-ledger val-ledger.jsonl
# ledger valid: 3 records
Each record is chained via SHA-256 of the previous record. Tampering is detectable.
Two subcommands. run is the default — if you omit the subcommand, val scenarios/ is equivalent to val run scenarios/.
val run <paths...> [options]Load scenarios from YAML files, execute steps, run L0 checks, optionally run L2 judge, and write a report.
| Option | Type | Default | Description |
|---|---|---|---|
--judge | "remote", "typesafe", "mock", "none" | "mock" (or "remote" if VAL_JUDGE_API_KEY or OPENAI_API_KEY is set) | Judge mode |
--judge-model | string | "gpt-4o" / VAL_JUDGE_MODEL for remote; "jev-latest" for TypeSafe | Model name for the selected judge |
--judge-base-url | string | "https://api.openai.com/v1" or VAL_JUDGE_BASE_URL env var | Base URL for an OpenAI-compatible endpoint |
--judge-api-key | string | VAL_JUDGE_API_KEY or OPENAI_API_KEY env var | API key for the remote judge |
--threshold | number (0–1] | 0.7 (or per-scenario judge.threshold) | Minimum overall score for L2 to pass |
--filter | string | — | Comma-separated tags; runs scenarios matching at least one requested tag and scenarios with no tags |
--bail | boolean | false | Stop after the first failing scenario |
--keep-workdir | boolean | false | Do not delete the temp working directory after the run |
--keep-ansi | boolean | false | Preserve ANSI escape sequences in captured output |
--report-dir | string | "./val-report" | Directory for the written report file |
--ledger | string (optional value) | "./val-ledger.jsonl" if flag present | Append a hash-chained ledger record. Use bare --ledger for default path, or --ledger <file> for a custom path. |
--parallel | positive integer | 1 | Maximum number of scenarios to run concurrently |
--freeze | boolean | false | For agent-mode scenarios, print executed steps as YAML for deterministic replay |
--target-ref | string | — | Evaluate against a branch, tag, or commit in a temporary detached target worktree |
For a ready-to-run Jev walkthrough, see demo/jev. It uses
TYPESAFE_API_KEY only at runtime and shows how to retain and compare TypeSafe and
remote judge reports.
val verify-ledger [file]Verify the hash-chain integrity of a ledger file.
| Argument | Description |
|---|---|
[file] | Path to the ledger file. Defaults to "./val-ledger.jsonl"; --ledger <file> may also supply the path. |
Exits 0 with ledger valid: N records on success, or 1 describing the mismatch on tamper detection.
Val is provider-agnostic. The judge adapter uses a plain OpenAI-compatible /v1/chat/completions endpoint — no SDK lock-in:
The repository includes a composite GitHub Action that installs Val, exposes passed and exit-code outputs, uploads the report and ledger even on failure, and then enforces the result:
- uses: stalewell/val@v1
with:
scenarios: scenarios/
judge: mock
target-ref: ${{ github.event.pull_request.head.sha }}
Pin the action to a full commit SHA in security-sensitive repositories. Until @stlw/val is published, set version to an installable package spec such as a Git URL.
| Input | Default | Description |
|---|---|---|
scenarios | required | Scenario file or directory passed to val run |
version | @stlw/val@latest | npm package spec to install |
judge | none | Judge mode or provider |
report-dir | val-report | Directory for JSON and Markdown reports |
ledger | val-ledger.jsonl | Path for the evidence ledger |
target-ref | — | Branch, tag, or commit to evaluate |
threshold | — | Passing score threshold |
parallel | — | Scenario concurrency |
bail | "false" | Stop launching scenarios after the first failure |
upload-artifacts | "true" | Upload report and ledger even when evaluation fails |
artifact-name | val-evidence | Uploaded evidence artifact name |
node-version | "20" | Node.js runtime used by Val |
| Output | Description |
|---|---|
exit-code | Val process exit code |
passed | Whether every scenario passed |
npm ci
npm run typecheck # tsc --noEmit
npm test # vitest run
npm run build # tsup → dist/
node dist/cli.js run scenarios/examples --judge mock
FAQs
External eval harness that proves your product does what it claims — deterministic L0 checks + LLM-judged L2 rubric scoring
The npm package @stlw/val receives a total of 174 weekly downloads. As such, @stlw/val popularity was classified as not popular.
We found that @stlw/val demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 2 open source maintainers collaborating on the project.

Company News
Socket is joining the OpenJS Security Stewardship Program to fund Node.js vulnerability research, maintainer remediation, and security releases.

Security News
Two compromised GitHub Actions were re-enabled with malicious tags intact, exposing thousands of downstream repositories to Mini Shai-Hulud.

Research
/Security News
A malicious Firefox extension fetches its payload after installation to evade detection, steal Google session cookies, and automate account takeover.