QualityMax QA MCP

Give a coding agent independent QA evidence before it declares a web change done: scan the page, inspect the UI, generate a focused Playwright repro, then review the execution result.
npx -y @qualitymax/qmax-mcp
The four local tools require no QualityMax account, API key, or hosted service.
Start with a useful result
Ask an MCP-enabled agent to scan the URL it changed, or run the local CLI:
npx -y @qualitymax/qmax-mcp scan https://example.com --format markdown
The report includes a graded summary, findings, concrete reproduction steps, and suggested fixes. One command exercises all four tools against a checked-in, dependency-free fixture — see the reproducible demo.

What one scan measures
All nine checks run off a single page load. Pass checks to run a subset.
console | JavaScript errors, warnings, and failed requests |
links | Broken and redirecting links, up to maxLinks |
accessibility | Missing alt text, unlabelled controls, nameless interactive elements, heading structure |
performance | Core Web Vitals: LCP, CLS, TTFB, FCP |
seo | Title and meta description |
security_headers | CSP, HSTS, X-Content-Type-Options, Referrer-Policy |
cookies | Missing Secure/HttpOnly/SameSite, third-party cookies, known trackers, tracking before consent |
mixed_content | HTTP subresources and form actions on an HTTPS page, split into browser-blocked active and passive |
weight | Transfer bytes, request count, render-blocking resources, oversized and uncompressed assets, third-party cost |
Names must match exactly. An unrecognised name is rejected with the supported list rather than skipped, so a typo cannot quietly turn a check off and still return a score.
scan_url also returns a metrics block with the measured vitals and the page-weight breakdown, including the slowest requests. Two limits are stated in that block rather than hidden: INP is not measured, because it needs real user interaction, and vitals come from one cold load on the scanning machine, not from field data. Set weightBudget to scan against your own performance budget.
What the report looks like
--format markdown opens with the shape of the result, so a human or an agent can see where the problems are before reading a single finding:
Grade: 🔴 F (0 / 100) · 17 issues found
░░░░░░░░░░░░░░░░░░░░░░░░ 0 / 100
| Console errors | ██░░░░░░░░ 1 | 🔴 high |
| Accessibility | ████████░░ 3 | 🔴 high |
| Security headers | ██████████ 4 | 🟠 medium |
| Cookies and trackers | █████░░░░░ 2 | 🟠 medium |
| Page weight | ██████████ 4 | 🟠 medium |
Each bar is scaled to the noisiest category in that run, so the tallest bar is the thing to fix first. When weight runs, the measurements section also attributes the bytes:
script ████████████████ 18 kB
document ██░░░░░░░░░░░░░░ 3 kB
stylesheet ██░░░░░░░░░░░░░░ 1 kB
image ░░░░░░░░░░░░░░░░ 417 B
Then every finding follows with its severity, a copy-pasteable reproduction, and a suggested fix. Use --format json for the same data as structured output.
What the agent can do
scan_url | Scan a URL for console, network, telemetry-SDK, link, accessibility, SEO, security-header, cookie/tracker, mixed-content, page-weight, and Core Web Vitals findings, optionally using a Playwright storage-state file. | It makes outbound requests, may read an explicitly selected workspace file containing credentials, and can write a screenshot. |
inspect_page | Return page structure and role/name locator candidates, optionally using a Playwright storage-state file for authenticated pages. | It makes outbound requests and may read an explicitly selected workspace file containing credentials. |
generate_playwright_repro | Write a deterministic, workspace-contained Playwright repro. | It writes below .qmax-mcp/repros; overwrites are explicit. |
run_playwright_test | Execute one local Playwright test and return structured status. | It executes code and writes controlled artifacts; by default it requires an accepted, digest-bound MCP human-approval elicitation. |
The local server does not require an account. Hosted proxy mode is a separate, opt-in connection for account-backed QualityMax capabilities; do not add it unless that capability is needed.
Scan or inspect an authenticated page
Both tools read a Playwright storage-state file. Producing that file is the
on-ramp, so here are the two shortest ways.
From a browser you log into yourself — no code, and it works against any
login, including SSO and MFA:
npx playwright codegen --save-storage=playwright/.auth/user.json https://example.com/login
Sign in in the window that opens, then close it. Playwright writes the state to
that path on exit.
From a login your test suite already automates — repeatable, and the one to
use in CI:
import { chromium } from 'playwright';
const browser = await chromium.launch();
const page = await browser.newPage();
await page.goto('https://example.com/login');
await page.getByLabel('Email').fill(process.env.APP_USER);
await page.getByLabel('Password').fill(process.env.APP_PASSWORD);
await page.getByRole('button', { name: 'Sign in' }).click();
await page.waitForURL('**/account');
await page.context().storageState({ path: 'playwright/.auth/user.json' });
await browser.close();
Read the credentials from the environment rather than writing them into the
script, and gitignore the output — playwright/.auth/ is the conventional
location and Playwright's own scaffolding already ignores it. A session expires,
so regenerate the file when scans start coming back as though logged out; a scan
of a protected page that suddenly reports a login form is the usual symptom.
Then pass the path relative to the active workspace:
{
"url": "https://example.com/account",
"storageStatePath": "playwright/.auth/user.json",
"acknowledgePrivateContent": true
}
scan_url and inspect_page load the state into a throwaway browser context
before the first navigation and require acknowledgePrivateContent:true as
explicit consent that the result may contain private page content. The path must
resolve to a regular file inside the workspace; absolute paths, traversal,
symlink escapes, and files over 10 MB are rejected. The state path and credential
values are not returned, but findings or inspected page structure may reflect
private account data. Treat both the state file and the tool result accordingly,
and use a dedicated state file containing only the credentials needed by the
inspected application. Playwright storage state covers cookies, local storage,
and optionally IndexedDB; it does not persist session storage.
Compare a scan against a previous one
scan_url is documented for use after a change, which makes it a comparison:
scan, change something, scan again, decide. Pass baseline — a previous scan
result, or a workspace-relative path to one — and the response carries a
delta reporting which findings are new, fixed, and unchanged, plus a
one-line verdict.
{
"url": "https://example.com/",
"baseline": "artifacts/scan-main.json"
}
Each finding carries a stable id derived from its category, message, and
URL or selector, so the same problem hashes the same across runs. Severity and
occurrence count are excluded from that identity on purpose: a finding that is
reclassified, or seen four times instead of three, is still the same finding.
In CI this is what makes a gate usable. Failing on findingCount > 0 stops
working the moment a page has one known-benign finding, whereas failing on
"nothing new since the last green run" keeps working:
qmax-mcp scan https://example.com/ --format json --out scan.json
qmax-mcp scan https://example.com/ --baseline artifacts/scan-main.json --fail-on-new
--fail-on-new exits 1 when the scan finds anything absent from the baseline,
and 2 when it was passed without --baseline. Persist the last main-branch
result as the baseline artifact.
Turn findings into tracker tickets
format: "issue" renders each finding as a self-contained Markdown block —
summary, numbered steps, expected result, actual result, environment — ready to
paste into a tracker without rewriting it as prose. minSeverity limits the
export to what is worth filing.
qmax-mcp scan https://example.com/ --format issue --min-severity medium
Blocks are separated by HTML comments, which render as nothing, so pasting one
does not carry the report's own structure into the ticket. A finding collapsed
from several occurrences files as one ticket that states the frequency, rather
than as several identical tickets.
Get ranked locators without an MCP client
inspect prints the same result as the inspect_page tool: every control with
a ready-to-paste Playwright locator ranked by how durable its source is, a
per-control stability verdict, and a page-level testability score.
qmax-mcp inspect https://example.com/
qmax-mcp inspect http://localhost:3000/ --allow-private-network --format json --out locators.json
The Markdown report leads with the testability verdict, then the locator table,
best handle first, with the caveats that make a fragile or none verdict
actionable. --format json returns the full structure for programmatic
consumers — a script choosing selectors for a generated spec, or a
selector-healing tool such as 9lives
picking the anchor to re-find a control by.
Add it to your coding agent
Copy a ready-made, no-credential configuration and the accompanying instruction file for your client:
The agent setup guide explains the expected approval surfaces and has a generic stdio configuration. The root AGENTS.md is the portable instruction: collect evidence, report unresolved failures, and request approval before mutating files or executing supplied code unless the server explicitly advertises unattended mode.
Unattended automation
For a trusted, isolated automation environment where no human can answer MCP
elicitations, start the server with the explicit --unattended flag:
{
"mcpServers": {
"qmax": {
"command": "npx",
"args": ["-y", "@qualitymax/qmax-mcp", "--unattended"]
}
}
}
For Codex TOML, use args = ["-y", "@qualitymax/qmax-mcp", "--unattended"].
This process-start opt-in authorizes every run_playwright_test call handled by
that server; it is intentionally not available as a tool argument or
environment variable. The server advertises the active mode to the agent,
prints an UNATTENDED startup warning, and returns
approval.mechanism: "unattended-cli-opt-in-v1" with each execution. The exact
test is still snapshotted and digest-checked, and the existing workspace,
environment, timeout, cancellation, and output controls remain active.
Adjacent QualityMax tools
The server tells a connected agent about three separate QualityMax tools that cover QA work these four tools do not. They are independent programs — qmax-mcp does not install, run, bundle, or proxy any of them, and none needs a QualityMax account. The agent is instructed to name one only when its trigger is present, once, and to leave the decision to run it with you.
| 9lives (MIT) |  | uv tool install 9lives, then 9l heal <spec> | A Playwright spec that used to pass is red after a change and the failure looks like drift. Heal the locator instead of weakening the assertion. |
| qualitymax-grader (Apache-2.0) |  | npx qualitymax-grader <spec> | A spec is about to be committed, or a suite is judged on test quality rather than on passing. Offline A-F grade, no model or network. |
| free-qa-skills (Apache-2.0) |  | install from skills.sh | The QA request is about a repository rather than a running URL, or the agent has no MCP server available. |
Together with the local tools they form one loop: scan_url finds the failure, generate_playwright_repro writes the spec, qualitymax-grader scores it before it lands, run_playwright_test executes it under the server's selected authorization mode, and 9lives heals it when a later change makes it drift.
Safety and honest limits
- Local scanning is networked. Private targets are denied by default;
allowPrivateNetwork: true is only deliberate caller-side consent for a narrow loopback target.
- Generated repros stay in a controlled workspace directory. Test runs use a minimal environment and controlled artifact directory.
- By default,
run_playwright_test uses MCP form elicitation before execution. The server displays the target, side effects, and a SHA-256 digest to the client, and runs only after the client returns an accepted human approval for that exact digest. Clients without form-elicitation support fail closed; a bare caller-supplied boolean is not accepted as proof. --unattended is the explicit process-level exception for isolated automation and permits supplied code to run with the local user's filesystem and network permissions without another human prompt.
- Read the full MCP safety contract and security threat model before publishing or enabling hosted capabilities.
Architecture

The launch comparison records dated, first-party capability references for TestSprite, BrowserStack, mabl, and Momentic. It is a factual boundary comparison, not a ranking.
Hosted proxy mode
The local tools are the open, local-first layer. Hosted QualityMax is an explicit proxy for workspace-backed project, test-case, script, and observability workflows:
QUALITYMAX_API_KEY="<your-api-key>" npx -y @qualitymax/qmax-mcp proxy
Only configure the proxy when a hosted-only capability is needed. The bearer credential is sent only to the pinned https://app.qualitymax.io/api/mcp/ endpoint; endpoint overrides and redirects are refused.
Support and responsible disclosure
Use GitHub Issues for non-sensitive usage and documentation support. Do not report vulnerabilities in a public issue; follow the repository security policy. The launch checklist includes the owner checks required before any public announcement.
Development
The runtime requires Node 22.13.0 or newer.
npm install
npx playwright install chromium
npm run check
npm run demo
npm run demo starts a dependency-free local fixture, prints a Markdown quality receipt by default, and leaves its generated repro and Playwright artifacts under .qmax-mcp/ for inspection. Use npm run demo -- --format json for a machine-readable receipt.
Candidate work after 0.4.0 — and the boundaries it has to respect — is recorded in the roadmap. It is a direction, not a delivery commitment; shipped changes are listed in the changelog.
Package and release metadata
Where the package is listed, where it deliberately is not, and what each channel accepts is recorded in the distribution channel inventory.
server.json is the canonical MCP Registry manifest. npm run validate:registry checks it against the official schema without publishing anything. Use npm run version:sync -- <semver> to move every public metadata surface — package.json, package-lock.json, server.json, smithery.yaml and src/metadata.ts — together; npm run check verifies they agree and fails on drift, and npm publish is restricted to the provenance-backed release workflow. The release workflow and rollback procedure are documented in the release runbook.
License
MIT — Copyright (c) 2026 QualityMax.