New:Microsoft Teams Notifications Are Now Available in Socket.Learn more →
Get Started

@stlw/val

Package Overview
Dependencies
Maintainers
2
Versions
3
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

@stlw/val

External eval harness that proves your product does what it claims — deterministic L0 checks + LLM-judged L2 rubric scoring

latest
Source
npmnpm
Version
0.1.2
Version published
Weekly downloads
177
1866.67%
Maintainers
2
Weekly downloads
 
Created
Source

Val

CI License: MIT

The eval harness that proves your product does what it claims.

Val drives a target application (CLI, HTTP), captures a transcript of real behavior, checks it deterministically (L0), and has an independent LLM judge grade it against the product's claims via a rubric (L2). Output doubles as evidence: a hash-chained, replayable record of "we claimed X; here is a machine-verified run demonstrating X."

See the project roadmap for authoritative implementation status and planned work.

npm install -g @stlw/val

How it works

Scenario (YAML)           Val harness              Judge (LLM)
┌──────────────┐       ┌──────────────┐       ┌──────────────┐
│ claim: "..."  │──────▶│  L0 checks   │──────▶│  rubric      │
│ steps: [...]  │       │  transcript  │       │  verdict     │
│ rubric: [...] │       │  captures    │       │  overall     │
└──────────────┘       └──────────────┘       └──────────────┘
                               │
                        ┌──────▼──────┐
                        │   report    │
                        │   ledger    │
                        └─────────────┘

L0 — Deterministic assertions on every step (exit codes, JSON shape, regex, negation). Free, runs on every commit.

L2 — An independent LLM judge scores the transcript against the product's claims via a rubric. Provider-agnostic — works with OpenAI, Anthropic, DeepSeek, Ollama, or any OpenAI-compatible endpoint.

Quick start

  • Write a scenario file (my-product.yaml):
id: my-product
product: my-app
claim: "The CLI exits cleanly with a helpful message on bad input."
steps:
  - name: store command
    run: my-cli store --content "hello"
    expect:
      - { exitCode: 0 }
      - { jsonPath: "$.id", exists: true }
    capture: { id1: "json:$.id" }
  - name: retrieve by id
    run: "my-cli retrieve ${captures.id1}"
    expect:
      - { exitCode: 0 }
      - { stdoutContains: "hello" }
  - name: bad input should fail
    run: my-cli store
    expect:
      - { not: { exitCode: 0 } }
      - { stderrContains: "--content" }
rubric:
  - id: clean-errors
    criterion: "Error messages are human-readable — no stack traces."
    weight: high
  - id: captures-work
    criterion: "Values captured from one step are correctly used in later steps."
    weight: medium
judge: { threshold: 0.7 }
  • Run it:
val run my-product.yaml --judge remote
  • Review the report:
cat val-report/report.md

Judge modes

ModeFlagDescription
Remote--judge remoteLive LLM via any OpenAI-compatible endpoint
TypeSafe--judge typesafeJev score judgments with confidence and probability evidence
Mock--judge mockDeterministic mock — always passes (use VAL_MOCK_JUDGE_FAIL=id to force failures)
None--judge noneSkip L2 entirely — L0 checks only
# OpenAI
val run scenarios/ --judge remote --judge-base-url https://api.openai.com/v1 --judge-api-key $OPENAI_API_KEY --judge-model gpt-4o

# Anthropic (Claude)
val run scenarios/ --judge remote --judge-base-url https://api.anthropic.com/v1 --judge-api-key $ANTHROPIC_API_KEY --judge-model claude-sonnet-4-20250514

# Local model via Ollama
val run scenarios/ --judge remote --judge-base-url http://localhost:11434/v1 --judge-model llama3

# Deterministic CI
val run scenarios/ --judge mock --ledger val-ledger.jsonl

# TypeSafe Jev (requires TYPESAFE_API_KEY)
val run demo/jev/claim.yaml --judge typesafe --judge-model jev-latest

Scenario reference

This README is a concise capability index. The scenario reference is the source for exact YAML semantics, defaults, template merge rules, scenario discovery, and complete examples.

Top-level fieldPurpose
idOptional stable scenario identifier (otherwise derived from the file path)
product, claimRequired product name and behavior claim
tagsLabels used by --filter
env, filesEnvironment variables and files created in the isolated scenario workdir
setup, teardownShell commands before and after the scenario
timeoutMsTotal scenario time budget
modesteps (default) or agent
goal, docs, maxStepsAgent-mode goal, product guidance, and step limit
stepsShell or HTTP interactions for step mode
rubric, judgeL2 criteria and per-scenario judge configuration
authHeaders injected into later HTTP steps
extends, importsReusable scenario-template references
Step fieldPurpose
name, descriptionOptional human-readable labels
runA shell command (exactly one of run or http)
httpAn HTTP request: method, url, optional headers and body
timeoutMsPer-step time limit
expectL0 assertions
captureNamed values extracted from the step result
allowFailureRecord a failed step and continue the scenario

Interpolation and captures

Use ${name} interpolation in supported scenario strings. The built-ins are ${workdir} (the isolated scenario directory), ${cwd} (the directory from which Val was invoked), and ${target} (the temporary target checkout when --target-ref is used). A capture named id is available to later values as ${captures.id}.

capture maps names to one of these expressions:

ExpressionCaptured value
json:<path>Value at a JSON path in stdout, for example json:$.id
regex:<pattern>First capture group, or whole match, from stdout or stderr
stdout:Trimmed stdout

Step types

Val supports two step adapters:

Shell — runs bash -c <command> against the target CLI:

steps:
  - run: my-cli store --actor alice --content "hello"
    expect: [{ exitCode: 0 }, { jsonPath: "$.id", exists: true }]
    capture: { id: "json:$.id" }

HTTP — drives REST APIs via fetch:

steps:
  - http:
      method: POST
      url: "http://localhost:3000/items"
      body: { name: "widget" }
    expect: [{ status: 201 }, { jsonPath: "$.id", exists: true }]
    capture: { itemId: "json:$.id" }

Expectations

L0 checks available on every step:

ExpectationApplies toDescription
{ exitCode: n }shellExit code equals n
{ status: n }httpHTTP status equals n
{ stdoutContains: s }bothStdout contains substring
{ stderrContains: s }bothStderr contains substring
{ outputContains: s }bothCombined output contains substring
{ outputMatches: re }bothCombined output matches regex
{ jsonPath: p, ... }bothStdout parsed as JSON; equals, exists: true, gte, lte
{ maxDurationMs: n }bothStep completed in ≤ n ms
{ fileExists: path }bothFile exists in the scenario workdir (path supports ${...} interpolation)
{ fileContains: { path, contains } }bothFile in workdir contains substring
{ not: <expectation> }bothNegates any expectation (nestable)

A timed-out step fails all its expectations. File paths resolve against the scenario workdir; paths escaping it fail the check.

Scenario timeout

Cap an entire scenario (setup + steps) with timeoutMs:

id: slow-suite
product: my-app
claim: "The full flow completes quickly."
timeoutMs: 30000
steps:
  - run: my-cli full-flow

Each command's timeout is capped to the remaining budget. On expiry the scenario fails with scenario timed out after Nms; teardown still runs.

Shared templates

Reuse config across scenarios with extends (single-inheritance) and imports (compose multiple blocks):

# profiles/base.yaml
product: my-app
tags: [smoke]
env:
  DB_URL: sqlite://test.db
setup:
  - my-cli provision --db-url "${DB_URL}"
# scenarios/login.yaml
id: login
claim: "Users can authenticate."
extends: ../profiles/base.yaml
imports:
  auth: ../profiles/auth.yaml
steps:
  - run: my-cli login --user alice

Merge rules: scalars override, objects deep-merge, arrays append (setup/teardown/tags). Steps come from the scenario file — templates provide env, files, setup, teardown, tags, auth, and judge config. Child rubric replaces parent rubric entirely. Template files live anywhere; they're loaded by relative path from the scenario file. Cycle detection prevents circular references.

Auth injection

Declare an auth block to auto-inject headers into HTTP steps after a capture is available:

auth:
  inject:
    headers:
      Authorization: "Bearer ${captures.token}"

After any step captures captures.token, subsequent HTTP steps automatically get the Authorization header. Step-level headers override injected headers. Auth is inactive before the capture is established (no error).

Agent mode

Let Val drive the product like a user, generating steps from a goal using the same LLM-backed judge endpoint:

id: explore
product: my-app
claim: "Users can create and retrieve items."
mode: agent
goal: "Create an item with content 'hello' and retrieve it by the returned ID."
docs: |
  Create: POST /items with { content } → { id, content }
  Get: GET /items/:id → { id, content }
maxSteps: 8
rubric:
  - id: crud-works
    criterion: "An item was created and successfully retrieved."
    weight: high

The agent receives shell + HTTP tools, observes outputs, and decides next steps. maxSteps (default 20) caps iterations. timeoutMs is checked only between agent iterations, so a single LLM call or action can exceed it; expiry does not currently mark the scenario timed out. After a successful run, --freeze outputs the executed steps as YAML for deterministic replay.

Parallel execution

Run scenarios concurrently with --parallel <n>:

val run scenarios/ --parallel 4 --bail

Results are reported in definition order. --bail stops launching new scenarios on first failure and waits for in-flight ones to drain.

Target references

Evaluate scenarios against another branch, tag, or commit of the current Git repository without switching the caller's checkout:

val run scenarios/ --target-ref main --judge mock

Val resolves the ref locally first. If it is missing and an origin remote exists, Val fetches only that requested ref. It then creates one temporary detached Git worktree for the run and exposes its path as ${target} and VAL_TARGET_DIR.

Scenarios remain responsible for preparing the target. Val does not guess the package manager or build command:

setup:
  - npm --prefix "${target}" ci
  - npm --prefix "${target}" run build
steps:
  - run: node "${target}/dist/cli.js" --version

Scenario commands still run in their isolated scenario workdirs. The target checkout is removed after the run, including when scenarios fail. --keep-workdir applies only to scenario workdirs.

In CI, the default shallow actions/checkout may not contain the requested ref. Val's narrow origin fallback handles branches and tags available on the remote; configure checkout history explicitly when testing arbitrary commit IDs.

Ledger

Every run can write a hash-chained, append-only record to val-ledger.jsonl:

val run scenarios/ --ledger
val verify-ledger val-ledger.jsonl
# ledger valid: 3 records

Each record is chained via SHA-256 of the previous record. Tampering is detectable.

CLI reference

Two subcommands. run is the default — if you omit the subcommand, val scenarios/ is equivalent to val run scenarios/.

val run <paths...> [options]

Load scenarios from YAML files, execute steps, run L0 checks, optionally run L2 judge, and write a report.

OptionTypeDefaultDescription
--judge"remote", "typesafe", "mock", "none""mock" (or "remote" if VAL_JUDGE_API_KEY or OPENAI_API_KEY is set)Judge mode
--judge-modelstring"gpt-4o" / VAL_JUDGE_MODEL for remote; "jev-latest" for TypeSafeModel name for the selected judge
--judge-base-urlstring"https://api.openai.com/v1" or VAL_JUDGE_BASE_URL env varBase URL for an OpenAI-compatible endpoint
--judge-api-keystringVAL_JUDGE_API_KEY or OPENAI_API_KEY env varAPI key for the remote judge
--thresholdnumber (0–1]0.7 (or per-scenario judge.threshold)Minimum overall score for L2 to pass
--filterstring—Comma-separated tags; runs scenarios matching at least one requested tag and scenarios with no tags
--bailbooleanfalseStop after the first failing scenario
--keep-workdirbooleanfalseDo not delete the temp working directory after the run
--keep-ansibooleanfalsePreserve ANSI escape sequences in captured output
--report-dirstring"./val-report"Directory for the written report file
--ledgerstring (optional value)"./val-ledger.jsonl" if flag presentAppend a hash-chained ledger record. Use bare --ledger for default path, or --ledger <file> for a custom path.
--parallelpositive integer1Maximum number of scenarios to run concurrently
--freezebooleanfalseFor agent-mode scenarios, print executed steps as YAML for deterministic replay
--target-refstring—Evaluate against a branch, tag, or commit in a temporary detached target worktree

For a ready-to-run Jev walkthrough, see demo/jev. It uses TYPESAFE_API_KEY only at runtime and shows how to retain and compare TypeSafe and remote judge reports.

val verify-ledger [file]

Verify the hash-chain integrity of a ledger file.

ArgumentDescription
[file]Path to the ledger file. Defaults to "./val-ledger.jsonl"; --ledger <file> may also supply the path.

Exits 0 with ledger valid: N records on success, or 1 describing the mismatch on tamper detection.

Works with

Val is provider-agnostic. The judge adapter uses a plain OpenAI-compatible /v1/chat/completions endpoint — no SDK lock-in:

  • OpenAI (GPT-4o, GPT-4o-mini)
  • Anthropic (Claude via beta endpoint)
  • DeepSeek (deepseek-chat)
  • Ollama (local models)
  • Groq, Together, Fireworks (any compatible endpoint)

CI

The repository includes a composite GitHub Action that installs Val, exposes passed and exit-code outputs, uploads the report and ledger even on failure, and then enforces the result:

- uses: stalewell/val@v1
  with:
    scenarios: scenarios/
    judge: mock
    target-ref: ${{ github.event.pull_request.head.sha }}

Pin the action to a full commit SHA in security-sensitive repositories. Until @stlw/val is published, set version to an installable package spec such as a Git URL.

GitHub Action reference

InputDefaultDescription
scenariosrequiredScenario file or directory passed to val run
version@stlw/val@latestnpm package spec to install
judgenoneJudge mode or provider
report-dirval-reportDirectory for JSON and Markdown reports
ledgerval-ledger.jsonlPath for the evidence ledger
target-ref—Branch, tag, or commit to evaluate
threshold—Passing score threshold
parallel—Scenario concurrency
bail"false"Stop launching scenarios after the first failure
upload-artifacts"true"Upload report and ledger even when evaluation fails
artifact-nameval-evidenceUploaded evidence artifact name
node-version"20"Node.js runtime used by Val
OutputDescription
exit-codeVal process exit code
passedWhether every scenario passed

Development

npm ci
npm run typecheck    # tsc --noEmit
npm test             # vitest run
npm run build        # tsup → dist/
node dist/cli.js run scenarios/examples --judge mock

Keywords

eval

FAQs

Package last updated on 20 Sep 2026

Related posts