jevnav
Page truth for browser agents — and decisions that replay, test and audit.

A coding agent working on a frontend codebase gets two things from jevnav:
- Page truth, not pixels. Structure, computed styles and controls come back
as facts, and
diff reports font-size 32px → 28px between a mockup and the
running app — the shape an agent can fix. No screenshots in the decision loop.
- Evidence, not confidence. Every action is a Jev decision with a calibrated
probability, risky ones are gated, and the whole run is a trace that
replay
re-checks offline in CI: a site change that breaks a recorded decision exits 1,
with no model call and no API key.
Selector-based tests break the moment a label changes, and LLM browser agents
are confident, unauditable and occasionally wrong. jevnav sits in between
(the longer version of this argument: docs/why.md):
- Jev picks the element. The candidate list of the current page is turned
into a choice question; the model answers with one element and a calibrated
probability. In loop mode (
jevnav go) one request also answers what to
do, whether the goal is already met and which context value to type.
- Every decision is recorded. The trace holds the candidates as the model
saw them, the choice, the probability and the cost — one JSONL file per run.
- Risky actions are gated.
p below the threshold, or an intent that looks
destructive, goes to a human instead of clicking.
replay is the regression test. Offline, no model call: re-resolve every
recorded decision against the page as it is now. A site change that breaks a
target fails CI; everything else is reported as drift, not noise.
Page truth for your agent
The facts a coding agent needs about a rendered page, without a screenshot:
outline(selector) for the region's structure (tags, headings, text, boxes),
styles(selector, props) for the computed values the browser resolved,
page_state() for the controls jevnav can see.
Mockup vs app, as facts instead of pixels — jevnav diff reports structure and
style differences and exits 1 on drift:
| element | property | mockup | app |
|---|---|---|---|
| h1 [Pricing] | font-size | 32px | 28px |
| button#cta [Start free] | border-radius | 8px | 4px |
A font-size 32px → 28px is something an agent can fix; a red pixel diff is
not. Once the app matches, pin the outcome (goal(..., success="<selector>"))
and replay --execute re-checks it in CI. Full walkthrough:
Matching a mockup to the app.
Install
uv tool install jevnav
playwright install chromium
Published on PyPI, listed in the
MCP registry
as io.github.dtduc-git/jevnav, and the replay Action is on the
GitHub Marketplace.
jevnav run needs a TypeSafe API key (TYPESAFE_API_KEY, or
~/.config/typesafe/apikey.txt). jevnav replay needs none — that is the point.
Quickstart — let Jev drive
jevnav go --goal "sign in with the demo account and open the pricing page" \
--start https://app.example.com/login \
--context email=demo@example.com --context password="${ACME_PASSWORD}" \
--success "#pricing.visible" \
--report goal.md
status: done — outcome verified against the page
steps: 5 — auto 4, review 0, blocked 0
One Jev request per step, and every step is gated and traced. The loop stops
when the model says the goal is done, when no listed element can make progress
(stuck), when the gate wants a human (review), when the page stops changing
(no_progress), or at --max-steps. --dry-run decides without acting.
done is a claim, not evidence. Pass --success <selector> and the claim
is checked against the page: verified, unverified (the selector is not
there — the run fails), or "not verified" when you passed no selector at all.
Quickstart — a scripted flow
id: acme-login
start: https://app.example.com/login
steps:
- intent: "Sign in to the existing account"
action: click
- intent: "Type the password"
action: fill
value: "${ACME_PASSWORD}"
- intent: "Submit the login form"
action: click
expect: "button[type=submit]"
jevnav run flows/acme-login/flow.yaml --report run.md
jevnav replay acme-login.trace.jsonl --report replay.md
jevnav diff new-ui.html http://localhost:3000
run walks the flow: extract candidates → ask Jev → gate → act → record.
replay re-checks the trace against the live site, with no model in the loop
(and with --execute it re-runs the recorded actions and verifies the recorded
--success selector, so a whole agent run becomes a CI test):
steps 3 verdicts: ok 3
Change Sign in to Log in on the site and the same replay reports:
[01] changed Sign in to the existing account
no element now has 'button|sign in' (was 'Sign in' / 'button')
Exit code 1, with the reason — that is the CI gate.
pytest — intents in an ordinary test
The jev fixture ships with the package, so a normal Playwright test gets Jev
decisions without changing how you write tests — and every test writes a trace
that replays in CI:
def test_sign_in(jev):
jev.goto("https://app.example.com/login")
jev.fill("the email address", "demo@example.com")
jev.fill("the password field", "${DEMO_PASSWORD}")
jev.click("the sign-in button")
jev.expect("#welcome")
DEMO_PASSWORD=... pytest --jev-trace-dir=traces
DEMO_PASSWORD=... jevnav replay --execute traces/test_sign_in.trace.jsonl
jev.expect is recorded on the trace, so the replay verifies the outcome as
well as re-running the actions. A review verdict fails the test before the
action runs, ${VAR} values are recorded by name only, and jev.page is the
real Playwright page for everything else. Runnable example, with a committed
trace anyone can replay: examples/pytest-interop/.
Games (the Doom shape)
A game has no candidate list to extract, so play takes the other shape: you
give it a JS state probe and a small action set, and Jev decides at a fixed rate
while the game keeps running — movement keys stay held between decisions, so the
last answer applies while the model thinks. Exactly how Jev plays Doom (fed
structured state as text, ~10 calls/second, no images).
jevnav play \
--goal "catch the green blocks, dodge the red ones" \
--url "examples/game/index.html?seed=7" \
--state-js examples/game/state.js \
--actions "left=ArrowLeft" --actions "right=ArrowRight" \
--rate 4 --seconds 60 --score-js "window.jevnavScore()" \
--ready-js "() => !!window.jevnavState" \
--report play.md
Measured on the bundled game (examples/game/), 60 seconds, three seeds, same
decision rate for both sides:
Jev latency was 312ms p50 from Vietnam, which is what caps the loop at ~3
decisions/second (TypeSafe's Doom demo ran ~10/s from a US network). Run
--policy random for your own control, and keep the trace: it is the evidence.
What works where: a DOM game (like the bundled one, or 2048) exposes state to
JS, so a probe is easy. A <canvas>/WebGL game — or Flash — has no state in the
DOM; it needs the game itself to expose one (Chocolate Doom WASM does, which is
how the browser Doom agents read it). jevnav still takes no screenshots and
makes no pixel decisions, deliberately: that is what keeps decisions replayable.
Your own Chrome (logins, cookies, extensions)
Three ways to get a browser:
jevnav go --goal "..."
jevnav go --goal "..." --user-data-dir ~/.cache/jevnav-profile --headed
jevnav go --goal "..." --cdp http://127.0.0.1:9222
--user-data-dir is a persistent Chromium profile: run once with --headed,
log in by hand, and every later run (headless or not) is already logged in.
Headful mode needs the full browser: playwright install chromium.
--cdp attaches to a Chrome you already have open — your session, your
extensions, the tab you are looking at. Start it with
--remote-debugging-port=9222 (or use chrome://inspect to find the port).
jevnav picks the last real page it finds, and never closes your browser.
Both flags work on run, go, replay and mcp. A trace records what was
decided, never which profile was used: cookies and profile paths never reach it.
Use the same freedom for authentication: pass credentials as context and let the
goal fill a login form when a stale cookie would be worse than a fresh login.
Gates
min_confidence: 0.9
loop_min_confidence: 0.5
risky:
- "\\b(delete|remove|purchase|pay)\\b"
intents:
"delete the *": { min_confidence: 0.99 }
truncated: review
Three verdicts, no ambiguity:
auto | confidence at or above the threshold, nothing risky — the action runs |
review | a human confirms first (low p, risky intent, truncated candidate list) |
blocked | no decision was possible (model answered none, or the call failed) |
n/a | the loop stopped itself (done) — no action to gate |
In loop mode the confidence threshold is lower on purpose. Measured
2026-09-21: correct loop decisions land at p 0.41–0.99 and wrong ones at
0.39–0.47, so p does not separate them. What keeps the loop safe is
deterministic: fill on a button is refused before it runs, a field with no
context value is blocked, two steps that change nothing stop the run, risky
patterns always go to review, and the outcome is verified against --success.
MCP — for other LLMs
jevnav runs its own browser and exposes it as an MCP server, so a coding agent
(Claude Code, Codex, Cursor, anything that speaks MCP over stdio) can drive a
page through Jev decisions instead of writing selectors:
jevnav mcp
No URL is needed: the agent opens pages itself with goto(url), so one server
serves every domain — a session can visit several sites, in several tabs. The
session writes jevnav-session.trace.jsonl in the client's working directory by
default (previous sessions are archived beside it, --no-trace opts out), so
every session is auditable without configuring anything.
--start <url> exists only as a convenience for a project-scoped config that
always begins on one page; put it in that project's config, not in your global
one. Same for the per-install choices: --browser, --user-data-dir (log in to
any number of sites once, in one profile), --cdp, --locale, --timezone.
Tools:
Deciding — the part no other browser MCP has:
browse(intent, action, value, min_confidence) | one step: Jev picks the element, the gate decides, and only auto acts |
goal(goal, context_json, max_steps, success) | drive the whole way: "sign in and open billing" — success verifies the outcome; returns done / stuck / review plus the verification |
goto(url) | open a page |
page_state() | URL, title and the shortlist jevnav can see |
summary() | this session: steps, auto/review/blocked, cost, latency |
Acting — everything else an agent needs:
screenshot(path, full_page, selector) | save a PNG for a human (never used by a decision) |
upload_files(paths, selector, intent) | set files, on a selector or an input Jev picks |
drag(source_selector, target_selector) | drag one element onto another |
resize(width, height) | change the viewport |
emulate(color_scheme, media, geolocation, offline, …) | emulate media, location and connectivity |
press_key(key, selector) | a key or combination ("Control+A"), optionally on an element |
fill_form(fields_json) | fill several fields in one call: {selector|intent, value, action} |
wait_for(text, selector, timeout_ms) | wait for something to appear |
scroll(direction, amount) | scroll the document |
tabs(), new_page(url), select_page(i), close_page(i) | work with tabs |
Inspecting — the agent's eyes (observation only, never traced):
console(limit, only_errors) | recent console messages and page errors |
network(limit, only_failed) | recent requests, with statuses |
network_detail(index, url_contains) | one request's headers and body |
dialogs() | alert/confirm/prompt, with the policy or rule that resolved them |
dialog_policy(action, match) | answer future dialogs: the default, or rules by message text |
read_js(expression) | evaluate JS in the page |
outline(selector, limit) | a page or region's structure (tags, headings, text, boxes) |
styles(selector, props, limit) | computed styles of the matching elements |
route(pattern, status, body, abort) / unroute(pattern) | stub or block requests (testing) |
trace_start() / trace_stop(path) | a Playwright trace zip for playwright show-trace |
perf_metrics(), heap_snapshot(path) | Chromium counters and a heap snapshot (best-effort: for real profiling use chrome-devtools) |
emulate(cpu_throttle, network_conditions, …) | CPU throttling and Slow-3G-style profiles (chromium, via CDP) |
lighthouse(url, categories) | Lighthouse scores, through npx (best-effort: needs node) |
Wire it into a client (this JSON shape is what Cursor, Claude Desktop and VS
Code use; Claude Code also accepts
claude mcp add --scope user jevnav -- uvx jevnav mcp):
{
"mcpServers": {
"jevnav": {
"command": "uvx",
"args": ["jevnav", "mcp"],
"env": { "TYPESAFE_API_KEY": "..." }
}
}
}
Why an agent would: it does not need its own Playwright MCP, it cannot click a
Delete by accident (review never executes; risky-action patterns ship for
nine languages, and extend them in gates.yaml), and its whole session is a
trace that jevnav replay --execute can re-run in CI. Cost is about
$0.00004 and 330ms per step; page_state and goto are free. Every tool
declares its MCP annotations — read-only, destructive, idempotent, open-world —
so a client can tell an observation from an action before calling it.
Dialogs: answered by rule, not parked
Playwright's sync API must answer a dialog inside its handler. Parking one so a
human can decide later blocks the renderer and the next call never returns
(measured on this codebase, then removed). So jevnav answers from a policy you
set in advance — dialog_policy("accept", match="delete") — and records every
dialog with the rule that fired, so the run stays auditable.
Which one to reach for
| who picks the element | a human writes selectors | the LLM, from a snapshot | Jev, with a calibrated probability |
| scope | the full test-authoring API | 29 tools, primitives + profiling | 33 tools, intent-level acting + observation |
| risky actions | whatever the test says | whatever the LLM says | never executed until a human says so (risky patterns cover English, Vietnamese, German, French, Spanish, Portuguese, Japanese, Chinese and Korean) |
| regression evidence | trace viewer, re-run the test | none | decision trace + offline replay that exits 1 |
| outcome assertion | expect(...) | none | --success selector, verified or reported unverified |
| engines | chromium, firefox, webkit | chromium | chromium, firefox, webkit (--browser) |
| CPU throttling / Slow-3G | ✅ | ✅ | ✅ (chromium, CDP) |
| request headers/body | ✅ | ✅ | ✅ network_detail |
| multi-field form fill | ✅ | ✅ fill_form | ✅ fill_form (selector or intent) |
| key combos | ✅ | ✅ press_key | ✅ press_key |
| per-step cost | 0 | one LLM turn per step (~38k chars of snapshot) | $0.00004 |
This is not a replacement argument: the three do different jobs, and running
more than one costs a line of config (jevnav's browser starts in ~20ms and is
lazy, so a second server is close to free). Playwright is the library you write
a test suite with — jevnav is built on it. chrome-devtools is what you reach for
to debug a page: screenshots, console, network, performance, all the raw detail
in the model's context, which is exactly right for debugging and exactly wrong
for driving. jevnav is the decision + evidence layer: an intent in, a gated
action out, a trace that replays offline. Reach for it when the same flow has to
keep working, and for chrome-devtools when you need to find out why it stopped.
When the gate says review
review is the tool refusing to guess — and it hands back what you need to
resolve it: the runners-up with their probabilities and a hint. Measured on
Wikipedia's main page (254 candidates):
goal("search Wikipedia for ...") → review at p=0.33 — the page has two
plausible ways to submit, so nothing was clicked.
- The caller re-reads the page and calls
browse("Click the Search button that submits the search form in the site header", "click", min_confidence=0.8)
→ auto at p=0.95, clicked for real.
goal(...) again → done, verified: true against .mw-search-results.
So: be specific, and if you know the page better than the model does, set your
own min_confidence — the confidence bar is the caller's call. Risky patterns,
role/value validation and the outcome check are not overridable.
The stdio path is tested end-to-end in CI: a real MCP client connects to a
jevnav mcp subprocess, lists the tools, calls goal and checks the browser
acted (tests/test_mcp_server.py, no network, fake Jev endpoint).
CI — the Action
The Action replays a recorded trace and fails when a site change breaks a
recorded decision. No model call, no API key, ~30 seconds:
- uses: dtduc-git/jevnav@v0
with:
trace: examples/local-demo/demo.trace.jsonl
execute: "true"
report: replay.md
Inputs: trace, report, execute, json, version (default latest from
PyPI, or local to run a checkout). Exit code 1 when a target changed, became
ambiguous, or a recorded --success selector is no longer visible. @v0 is a
floating tag; pin @v0.1.1 if you prefer.
How it works
- Candidates are a shortlist, not the page. Visible interactive elements
ordered by how likely a human would act on them — in-viewport first, form
controls before buttons before links — capped at 120 (
--max-candidates, the
API's hard cap is 254). Measured on Hacker News (199 elements → 40): same
accuracy, 2.8× faster on a cold decision and 3.8× fewer input tokens.
Each carries role, accessible name, type, href, placeholder and a scope
(nearest legend/heading) so three "Email" fields stay distinguishable.
- Fingerprint. An element's identity is
role|name (whitespace- and
case-normalized). Traces store the fingerprint of every candidate as it was
shown to the model, so replay never re-derives identity with new code.
- Decisions. One choice question per step: the option map is the candidate
list, plus
none. The decision is recorded with probabilities, usage and cost.
- Replay. Re-extract the page, compare fingerprints.
--normalize REGEX
relaxes matching for known churn (a counter like "Cart (3)" → "Cart (4)") on
both sides, opt-in, so strict is still the default. ok (found),
moved (found elsewhere on the page), changed (gone), ambiguous (now
duplicated), error. changed, ambiguous and error fail; moved and
drift counts are reported.
- Actions.
click, fill, select, check, hover, press, none.
Replay re-runs actions only with --execute, and resolves them by
fingerprint — never by position — so a shifted page cannot click the wrong
thing.
Matching a mockup to the app
jevnav diff new-ui.html http://localhost:3000 --report ui-diff.md
- mockup: `new-ui.html` — 'Pricing (new UX)', 7 elements
- app: `http://localhost:3000` — 'Pricing', 4 elements
- differences: **5** structure, **4** style
## Structure (`body`)
| kind | element | detail |
|---|---|---|
| missing | p 'Three plans for every team.' | not on the other page |
| missing | section 'Enterprise Talk to sales' | not on the other page |
| missing | button 'Talk to sales' | not on the other page |
| new | button#extra 'Book a demo' | only on the other page |
| moved | h1 'Pricing' | x+0 y+0 w+0 h-5px |
## Styles (`h1,#cta`)
| element | property | mockup | app |
|---|---|---|---|
| h1 [Pricing] | font-size | 32px | 28px |
| button#cta [Start free] | border-radius | 8px | 4px |
Exit code 1 when anything differs, 0 when the pages match — so the same command
works as a CI check that the app has not drifted from the design. Structure is
matched by tag + the element's own text (self-closing containers are not
"changed" when a child disappears), boxes are compared with a 4px tolerance
(--tolerance), and fractional pixel values are rounded so layout noise does
not read as a change.
The loop for "here is a new UX/UI, update the codebase": the coding agent opens
the mockup and the running app with jevnav, reads the facts instead of
guessing — outline("main") for the structure, styles("#hero", ["font-size", "gap"]) for the computed values, page_state for the controls, screenshot for
the human — diffs the two, edits the code itself (that part is the coding agent,
not jevnav), then re-reads the app to confirm. goal("...", success="<selector>")
pins the result so the fix can be replayed in CI later.
jevnav reports; it does not edit your repository, and it does not compare pixels.
Architecture

docs/architecture.html is the interactive version (pan, zoom, themes, three
guided views: one decision, evidence and replay, the other loops); the spec it
was built from is docs/architecture.archify.json. In one line: the caller
gives an intent, jevnav reads a ranked shortlist from the browser, Jev picks
with a calibrated probability, the gate decides whether that may run unattended,
the action goes back through the DOM, and every step lands in a trace that
replay re-resolves offline.
Why the loop is cheaper: two sequences

Same task, different anatomy. With chrome-devtools-mcp the LLM is the eyes:
every step it reads a ~38k-character accessibility snapshot into its own context
(~10k tokens on a frontier model), decides the element, clicks, and pays for a
full turn again on the next step. With jevnav the LLM asks once (goal), and
each step is a ~330ms, $0.00004 question to Jev over a ≤120-candidate shortlist
that never enters the LLM's context — with a gate in between and a trace written
as it goes. An early one-run sample with the same LLM (deepseek-v4.1-flash via
opencode) is in the git history; do not lean on it — n=1 per server, and its
loudest number came from a robot-policy 403, not from architecture. The
deterministic claim is replay, and it needs no benchmark to defend.
Interactive versions of both sequences: docs/seq-chrome-devtools.html,
docs/seq-jevnav.html.
Benchmarks
On a driving task set (a local ops console: sign-in, a form inside a shadow
root, a table row action, an iframe invoice), same cheap LLM for both servers,
n=2 per task: jevnav 8/8 tasks, chrome-devtools-mcp 6/8 — and the two
failures were model flakiness, not capability (a manual rerun finished with the
right answer through the shadow root). chrome-devtools was 2.4x faster
end-to-end (19.3s vs 45.6s mean) with fewer calls. That is the honest
correction to any "faster" claim: jevnav's advantage is decision cost and
evidence, not wall clock on small pages. Full method and caveats:
research/driving-benchmark.md.
Two more numbers, and only one of them is a comparison.
Deterministic, and the one to hold jevnav to: replay is offline, needs no
API key, and exits 1 when a recorded decision no longer resolves. There is no
sampling error in that; run it on your own traces.
Tool-level, and weaker by nature — benchmarks/mcp-compare.py, same machine,
one task, against chrome-devtools-mcp:
| MCP ready | 22ms (lazy browser) | 491ms |
| observation the agent must read | 4.8k chars | 38.3k chars |
| tool calls for the task | 2 | 4 |
| decision cost (real / modelled) | $0.0008 | $0.057 |
| outcome verified against the page | yes (--success selector) | no such notion |
An earlier run with the same LLM (deepseek-v4.1-flash via opencode) is on
record in the git history, but do not lean on it: n=1 per server, one model, two
tasks, and the loudest number (a Wikipedia search where chrome-devtools took 83s
and hit HTTP 403) is a robot-policy artifact, not an architectural difference.
The honest version is the table above — what the caller pays per step and how
much of the page lands in the model's context — and even that says nothing about
how the two behave across many sites. What jevnav claims is narrower and provable
on your own pages: a decision at or above the gate is safe to run, and the run
replays.
Measured
The goal loop, measured on 2026-09-21 (4 goals × 2 wordings × real Jev, local
fixture: sign in, open pricing, sign in then pricing, an impossible goal):
8/8 goals correct, including the impossible one (stuck), $0.00004 per
step, p50 314ms per step. One real run — sign in then open pricing — took 5
steps, $0.000214, and replayed offline with --execute: 5/5 targets resolved,
outcome verified.
On real pages (research/browser-element-selection.md, 30 hand-labelled cases
across 8 public sites, one decision each, model jev-1.13.0):
- 41/41 scored cases correct; 30 ran at
p >= 0.9 and all 30 were right.
- 365ms p50, $0.000153 per decision.
- 71 cases are written, but only 41 scored: the harness refuses labels whose
selector matches zero or several visible elements, and 30 of mine did. Small n,
single annotator, well-built pages: a direction, not a proof. The cases, the
runner and the excluded-case log are all in the repo.
The build-time element-decision spike (44 decisions: local fixtures, Hacker News,
PyPI, Wikipedia):
- 44/44 decisions correct; 28/28 at
p ≥ 0.9 (the auto gate).
- Replay caught 4/4 injected DOM changes with 0 false alarms on the
unchanged pages.
- Latency p50 334ms, p95 834ms; $0.000053 per decision.
- Asked for an element that does not exist, Jev answered
none at p=1.0
and p=0.92 instead of inventing one.
Small sample, self-graded ground truth, easy intents — treat these as direction,
not proof. replay is the number that matters in CI, and it is deterministic.
Non-goals
- No planner and no agent loop — you (or your agent) decide what to do; jevnav
decides where and records why.
- No screenshots in the decision loop, no text generation (
fill takes the text
from your flow or your environment).
- No iframes, shadow DOM, canvas or file pickers in v0.1 — long tail, tracked as
issues rather than half-supported.
- No SaaS, no hosted runner, no telemetry. Local-first: nothing leaves the
machine except the question sent to your configured Jev endpoint.
Privacy
Traces contain page URLs, element names and your actions — never screenshots.
Loop mode also sends a short digest of the page's visible text (it is how the
model judges whether the goal is done) and the current value of form fields
(passwords masked) — that is what any browser agent has to observe. Scripted
flows send neither. Literal values from the flow are recorded (they are already in
your repo); ${ENV} values are recorded as the variable name only. Add
*.trace.jsonl to your project's .gitignore (jevnav's own repo does), and
audit a trace before sharing it.
Suite
jevnav is the browser piece of a verification stack: mcplint
(MCP configs), harnessguard (agent
harnesses), jevassert +
jev-packs (calibrated decision
packs), and jev-table.
License
Apache-2.0.