page-scanner
Capture a whole web page from the command line, as a PDF whose text is still text.
npx @page-scanner/cli install
npx @page-scanner/cli scan --url https://en.wikipedia.org/wiki/PDF --out ./pdf.pdf
The capture is done by the Page Scanner Chrome extension, in a Chrome you are already signed
in to. This package is the other half: it listens on loopback for that extension to connect, and
turns scan into a file on your disk. A page behind a login is a page you are logged into, and no
second browser is started.
It is also the engine behind @page-scanner/mcp, which is the same
operations as MCP tools for an agent.
Requirements
- Node.js 24 or newer.
- Google Chrome, or another Chromium browser, with the Page Scanner extension 1.3.0 or newer
installed. Earlier versions have no helper to connect through, no
--hide and no design.
Install
Two ways, in the order most people want them:
npx @page-scanner/cli --help
npm install -g @page-scanner/cli
Setting up
Once per machine:
page-scanner install
It sets up Page Scanner's helper for every Chromium browser it finds (Chrome, Chromium, Edge,
Brave, Arc, Vivaldi, and Chrome's Beta, Dev and Canary), then says which. Open the Page Scanner
settings in Chrome (the gear in the editor toolbar), go to Local Agents, press Connect,
and allow Chrome's prompt. That is all: there is no port and no token to copy.
What it writes: a Native Messaging host manifest in each browser's NativeMessagingHosts
directory (on Windows, a registry value under HKCU), and the helper itself in
~/.page-scanner/native-host/. Chrome starts the helper for the Page Scanner extension and no
other; the helper reads the pairing in ~/.page-scanner/config.json, which install creates if
there is none, and connects the extension to the daemon below. page-scanner uninstall takes it
all out again and leaves the pairing.
- An unpacked build loaded in developer mode is found by itself:
install reads each profile's
list of extensions (and nothing else in it) for a copy of Page Scanner loaded from disk, and
allows it too. Chrome writes a newly loaded one down about 10 seconds after loading it, so run
install again if it was that quick. --extension-id <id> adds an id by hand.
--browser-dir <dir> sets up only that user data directory, for a profile started with
--user-data-dir.
- Run it again after installing another browser, or after removing the Node it names: the helper
is started with the Node that ran
install, because Chrome starts it without your shell's
PATH. status says when that Node is gone.
--node <path> names another Node for the helper. Node reports its own path with symlinks
resolved, so a Homebrew Node is named by its versioned Cellar path, which the next
brew upgrade removes; --node /opt/homebrew/bin/node follows every upgrade instead.
Chrome runs a separate copy of the extension in every profile, so your work profile and your
personal profile are two different browsers here. Naming them in the settings is how you tell
them apart later.
Pairing with a port and a token
The older way, for a machine where the helper cannot be installed, and still working for every
browser paired this way:
page-scanner pair
It prints a port and a token and saves them to ~/.page-scanner/config.json. In the settings'
Local Agents, set Connect through to A port and a token, paste both, name the browser,
and press Connect. pair waits up to two minutes and tells you when the browser arrives.
page-scanner pair --rotate issues a new token and invalidates the old one. Every browser paired
by hand has to be re-paired after that; the helper reads the new one by itself.
Updating
npx @page-scanner/cli asks npm for the newest version each time it runs, and installs it when it
is newer, so there is nothing to update. A global install stays on the version it installed, and
while it is there npx runs that copy too:
npm install -g @page-scanner/cli@latest
The next command that finds the old daemon still running shuts it down and starts its own. The
pairing in ~/.page-scanner/config.json carries over, and the extension reconnects by itself, so
nothing needs pairing again. The helper install copied does not need updating either: it only
carries frames, and it reaches whichever daemon is running.
If you also use the MCP server, keep the two on the same version. A command replaces a daemon of
any other version, or of another build of the same version (which happens when the CLI is built
from source), so two versions on one machine keep replacing each other's daemon, and the browser
drops its connection each time.
Commands
page-scanner install [--extension-id <id>]... [--browser-dir <dir>]... [--node <path>] [--json]
page-scanner uninstall [--browser-dir <dir>]... [--json]
page-scanner pair [--port <n>] [--rotate] [--wait <s>=120] [--no-wait] [--json]
page-scanner status [--json]
page-scanner browsers [--json]
page-scanner tabs [--browser <id|label>] [--wait <s>=30] [--json]
page-scanner scan (--url <u>... | --urls <file|-> | --tab <id>) [--window <id>]
[--name <template>] [--markdown beside|only]
[--main-content] [--max-chars <n>]
[--links inline|references|text] [--pictures omit|alt]
[--redact]
[--hide <kinds>|all|none]
[--slices] [--slice-side <px>=1568] [--picture-files]
[--accessibility] [--tables]
[--highlight <quote>]... [--highlight-from <file|->]
[--browser <id|label>]
[--format pdf|png|jpeg=pdf] [--page-size auto|a4|letter|phone=a4]
[--quality <0-1>] [--video frame|blank]
[--scheme auto|light|dark]
[--page-width window|a4|letter|phone] [--open-editor]
[--out <file|dir>] [--wait <s>=30] [--timeout <s>=120] [--json]
page-scanner diff <old> <new> [--section <heading>] [--out <file>] [--json]
page-scanner verify <file> [--record <file.integrity.json>] [--json]
page-scanner check <file.md> [<quote>...] [--from <file|->] [--loose] [--json]
page-scanner design (--url <u>... | --urls <file|-> | --tab <id> |
--crawl <u> [--max-pages <n>=10] [--depth <n>=2])
[--components] [--window <id>]
[--out <dir>] [--min-uses <n>=2] [--browser <id|label>]
[--wait <s>=30] [--timeout <s>=120] [--json]
page-scanner serve [--daemon] [--idle <min>]
page-scanner stop [--json]
page-scanner --version | --help
--browser takes a browserId or a label, matched without regard to case, and is only needed when
more than one browser is connected.
--out is a file when it ends in a 2 to 5 character extension and is not an existing directory,
and a directory otherwise, in which case the browser's suggested filename is used inside it.
Parent directories are created. A relative path is relative to your working directory, not the
daemon's, which is why the daemon hands the bytes back rather than writing them itself.
--page-width lays the page out at the width of a sheet before capturing it, the way a narrow
window would, so the PDF prints at 1:1 rather than being scaled down. window, the default, is
the width the browser has the page at. A 1280 px window on A4 is scaled to 55 %, which puts 16 px
body text at 6.5 pt; --page-width a4 puts it at 12. --page-width phone lays it out 390 px wide, as a phone shows it,
and --page-size phone cuts a PDF into phone screens, 390 × 844 px, with no margin.
--scheme picks which of a page's two themes to capture, for a page that has both. auto, the
default, is whichever the browser is showing; light and dark force one. It applies to the PDF
as well as the preview, which is the point: Chrome prints every page in the light scheme unless it
is told otherwise, so without this a dark page produced a white PDF.
--page-size decides what a PDF is laid onto. a4 and letter slice the capture across printable
sheets with a half-inch margin; auto is one page the exact size of the capture, which is not
printable and which Acrobat clamps past 200 inches.
--url more than once, or --urls <file> (one address to a line, # for a comment, - for
standard input), scans a list, one page at a time, into --out, which is then a directory. Each
path is printed as its file is written; a page that fails is reported on stderr and the rest carry
on, and the exit code is 1 if any failed. Up to 1,000 addresses a run.
--name names the file inside --out: {n} (the page's place in the list), {host}, {name}
(the browser's suggested name), {date}, {time} (when the run started) and {ext}, which is
added when left out. A / makes a subdirectory:
page-scanner scan --urls reading.txt --out ~/Captures/ --name '{date}/{n}-{host}'
--markdown beside also writes the page's text as a .md next to the file, read from the page
itself rather than from the PDF, and prints its path on a second line; --markdown only writes the
.md alone. With --json, the result carries the title, address, capture time, language and
headings as page. Pictures are left out of the text.
For an agent's context window, --main-content writes only the main content the extension found
(the whole page when it found none), --max-chars <n> keeps whole blocks up to that many
characters and reports in page.cut how long the whole was and how many blocks it left out,
--links references lists each address once at the end and --links text drops them, and
--pictures alt leaves a placeholder with each picture's alt text. Each heading in page.headings
has offset and chars, where its section starts in the text and how long it is. Dropping the
addresses is what shortens a link-heavy page most: Korean Wikipedia's "PDF" went from 85,657
characters to 37,134.
--redact replaces what Sensitive Text's rules find in the text before it leaves the browser:
every email address, phone number, card and IBAN, ID or tax number, IP address, key and token,
and the patterns you saved in the extension's Sensitive Text settings, each with a numbered
placeholder such as ⟦EMAIL 1⟧, the same value always the same placeholder. page.redacted counts
what was replaced by kind, never the values. It is for a signed-in page whose text goes on to a
cloud model; only the Markdown is redacted, so use it with --markdown only, since a PDF beside
it holds the page as it was. An extension older than this refuses the scan rather than send the
values.
--hide ads,consent,chat,overlays (or all, or none) hides ads, cookie and consent banners,
chat widgets and pop-ups before the capture and puts them back after; without it the extension's
own Clean up before capture settings decide. What was hidden is said on stderr and counted in
--json as hidden.
--slices also writes the capture as PNG slices for a vision model, top to bottom, into a folder
beside the file (page.pdf gives page.slices/slice-01.png and on). Each is at most
--slice-side pixels on a side, 1568 by default, the size past which Claude shrinks a picture,
and repeats the last 48 px of the one before, so a line on a seam is whole in one. A screenshot of
a whole long page is shrunk until its text cannot be read; a slice is not. --picture-files
writes the page's pictures and canvases (at least 120 by 40 px, up to 30 and 50) into
page.pictures/, cut from the capture where they were drawn. --json lists each file with the
band or box of the page it shows, in CSS px, as slices and pictures; slicesCut is set when a
page runs past 60 slices.
--accessibility checks the page against WCAG 2.0, 2.1 and 2.2, levels A and AA, with axe-core,
and writes page.accessibility.pdf, the capture with each failing element boxed and numbered and
the rules listed after it, and page.accessibility.json, every rule and element with its box on
the page in CSS px. --json has accessibility with both paths and the counts; stderr says how
many elements failed, by impact. An automated check finds some of what fails WCAG, not all of it.
--tables writes the page's tables as page.tables.json, { "tables": [...] }: each with
rows of cell text (every row as wide as the widest, a cell's words on one line, an icon by its
name), header (whether the first row is the page's header row) and box, where it was on the
page in CSS px. A <table>, an ARIA table (role="table" or "grid") and a CSS grid drawn as one
all count; a virtualized grid has only the rows it had drawn. --json has tables with the path
and each table's rows and columns; stderr says how many. With --markdown they cover what the text
covers, --main-content and --redact included.
--highlight "<quote>", once a passage, or --highlight-from <file> (one a line, or a JSON
array; - reads standard input), marks passages quoted from the page where the page has them as
written, found the way check finds them, and a PDF lists each under a Quoted passages
bookmark. --json says which were found as highlighted; one not found is named on stderr and not
marked. It is for one page, not a list.
When the page declares how to cite it (citation tags, schema.org data, Dublin Core, OpenGraph),
--json has it as citation: the kind of work, and the title, authors, published date,
container (journal or site), volume, issue, pages, doi and the rest it gave, as it wrote
them and unchecked, with from, where they came from. The file itself carries the same. A page the
browser had machine-translated before the scan has translated ({ "by": "chrome" }, or
edge), and the file says it is a translation.
page-scanner diff <old> <new> says what changed between two captures, with no browser: two
Markdown files (or two PDFs with the .md beside them) a passage at a time, printed as diff -u
prints it, or two PNGs pixel by pixel, writing the newer one with the changed regions outlined. It
exits 0 when nothing changed and 1 when something did. --section "Pricing > Pro" compares only
the section under that heading, so a banner or a rail of latest posts around it does not count.
page-scanner verify <file> checks a file against the integrity record the extension wrote beside
it: the SHA-256 and size, and the signature when there is one. It exits 0 when both hold and 1 when
either does not.
page-scanner check <file.md> "<quote>"... says whether each quote or value an agent is about to
use occurs in a capture's Markdown as written, and where: the line and the section (the headings
above it, or front matter). The Markdown's own markup is undone first (escapes, emphasis, link
syntax, list and quote prefixes, table pipes) and line wrapping is ignored; a quote never runs
across two blocks, and a word or number has to stand apart (4.99 is not in $14.99). For a
missing quote it prints where the capture parts from it. A quote that matches only once curly
quotes, dashes and case are folded is reported apart, and --loose accepts it. Quotes come after
the file or from --from (one a line, or a JSON array; - reads standard input). No browser and
no model.
page-scanner design --url <address> reads a page's design as the browser draws
it and writes six files into --out, or design-<host> in the working directory. A design
system is spread over a site, so --url can be given more than once, or --urls <file|-> can
list up to 50 pages of it: they are read one after another and written as one design, and
audit.md adds the values only one of the pages uses. --crawl <address> finds the pages
instead: it follows the page's links to its own site, breadth first, reading one page of each
kind (one blog post, not all, but each section that has pages below it), up to --max-pages (10) and --depth links away (2). It
follows <a href> only, honors robots.txt and nofollow, and never opens an address that
acts, such as a log-out link, since the tabs are your own signed-in Chrome. A page that a
redirect took to another site, or to a page already read, is left out, and one the server
answered with an error (a 403 or a 404) is not read in. The six files: tokens.json
(W3C Design Tokens format), tokens.css, tailwind.preset.js, audit.md (values nearly equal,
probably meant as one), contrast.md (WCAG 2 contrast for every text color over its background)
and extract.json (the measurements, for an agent). The page is read in its light and its dark
scheme; an address opened for it is reloaded under each, a --tab you have open is not. Custom
properties keep their names, except a declared color nothing is drawn in, which is left out and
counted. Several names of one value, or of two colors nobody could tell apart, are one token under
the most used name, the others listed as its aliases, and a declared length counts only what its
name says it is for (--radius-* radii, --space-* spacing, --text-* font sizes); other values become tokens when at least --min-uses elements use
them, numbered by use. The folder's path goes to stdout and a summary to stderr. The
extension has to be one that says it can (an older one is refused with a hint to update).
--components finds the components as well: the buttons, fields and repeated boxes on the pages,
each with its variants (the instances that look alike) and what changes on hover and focus. Each
page is also scanned, and three more files are written: components.json, components.md, and
catalog.pdf, one section per component with each variant cut from the page as vector artwork
and its measured values beside it.
There is no scheduler here: cron, launchd and Task Scheduler run the command at a time. Start the
daemon at login with page-scanner serve --daemon --idle 0 so the extension is connected when the
schedule fires. The documentation has an example for each.
--wait 0 means do not wait at all. Chrome retires the extension's service worker after about
thirty seconds of silence, so the default wait exists to cover the reconnect that follows.
What a command prints
- stdout carries the answer and nothing else. For
scan that is one line, the absolute path
of the file it wrote. For tabs and browsers it is a column-aligned table with no separator
row, so awk can read it.
- stderr carries anything addressed to a person: progress, warnings, the reason something
failed.
--json puts exactly one JSON document on stdout, for success and for failure alike, and
leaves stderr empty.
{
"ok": true,
"path": "/Users/you/pdf.pdf",
"width": 1280,
"height": 4200,
"mode": "vector",
"selectableText": true,
"truncated": null,
"browserId": "b-9f2c41",
"fileName": "PDF - Wikipedia.pdf"
}
{ "ok": false, "code": 3, "error": "NO_BROWSER", "message": "No browser is connected." }
truncated is null rather than absent when the capture was whole, so a reader can see the
question was asked. When it is not null it carries what the page measured, what was captured, and
a sentence naming the gap.
Exit codes
| 0 | success |
| 1 | the browser was reached and the work failed |
| 2 | the arguments were wrong |
| 3 | no usable browser: none connected, several connected, or the one named is not |
| 4 | not paired |
| 5 | the daemon would not start |
diff uses the codes the way diff itself does: 0 when the two captures are the same, 1 when
something changed, and 2 for anything it cannot compare. check does the same: 0 when every quote is
there, 1 when one is not, 2 for a missing file or no quotes.
The daemon
The extension is the WebSocket client: nothing outside Chrome can open a connection into a
service worker, so something has to be listening before Chrome can dial in. That something is a
background daemon, started on demand by the first command that needs it.
page-scanner scan ──HTTP /rpc──┐ ┌──< Chrome extension (paired by hand)
page-scanner-mcp ──HTTP /rpc──┼──> page-scanner daemon ──ws :45711
Claude Desktop ─HTTP /rpc───┘ (127.0.0.1) └──< helper ══stdio══ Chrome extension
The helper is what Chrome starts through Native Messaging. It dials 45711 the way the extension
would, with the pairing's token, and waits for the daemon if none is running yet; so the daemon
is still started by whatever wants a scan, and the browser arrives within a couple of seconds.
Two ports, deliberately. 45711 is the one the extension dials, and it refuses anything whose
Origin is not a chrome-extension://, which is exactly what a local CLI process is. The RPC port
is ephemeral, carries a bearer secret regenerated on every run, and is what the commands talk to.
page-scanner serve runs it in the foreground, which is the way to see why it will not start.
page-scanner stop stops it. It exits on its own after 15 minutes with no browser connected and
nothing calling (PAGE_SCANNER_IDLE_MINUTES, 0 for never), and never while a browser is
attached: the open socket is what keeps Chrome's service worker alive, and dropping it would make
the next command wait for a reconnect.
Files
Everything lives in ~/.page-scanner, or in $PAGE_SCANNER_HOME if that is set.
config.json | the pairing: { port, token }, mode 0600 |
daemon.json | the running daemon: pid, RPC port, secret, version. Mode 0600 |
daemon.log | stdout and stderr of a daemon started in the background |
native-host/ | the helper Chrome starts, and the script that starts it |
Windows: those 0600 modes are a no-op. NTFS does not implement POSIX permissions, so both files
are readable by anything running as you. That is usually fine on a personal machine. On a shared
one, restrict the directory yourself:
icacls %USERPROFILE%\.page-scanner /inheritance:r /grant:r %USERNAME%:F
Security
The token is the only thing between a local process and a scan of any tab you have open, because a
loopback port is reachable by everything running as you. So:
- The bridge binds
127.0.0.1 only, never 0.0.0.0.
- Every connection must present a
chrome-extension:// Origin and the shared token in its
first frame. The Origin check alone is not enough: a local process can forge that header.
- The extension does no port discovery. Loopback
fetch is blocked for an extension origin by
Local Network Access, so probing for a port would need a host permission the extension does not
have. The helper reads the port from the pairing file instead.
- Chrome starts the helper only for the extension ids its manifest names: the two store builds'
(the Chrome Web Store's and Edge Add-ons'), the unpacked copies of Page Scanner
install found, and any --extension-id you added. A store
extension of another id is never added by the search, whatever it is called. The helper holds no secret of its own; it reads the pairing
from a file only you can read, the same file pair writes.
- The token is printed by
pair, on stdout, and never appears in any other output or in the log.
- The extension refuses tabs Chrome will not let it touch:
chrome:// pages, the Web Store, and
the PDF viewer.
Using it from Node
import { scan, scanMany, listTabs, checkQuotesInFile } from '@page-scanner/cli';
const { tabs } = await listTabs();
const result = await scan({ tabId: tabs[0].tabId, out: './out/', format: 'pdf' });
console.log(result.path, result.selectableText);
const batch = await scanMany(['https://example.com/a', 'https://example.com/b'], { out: './out/' });
const { allFound, quotes } = checkQuotesInFile('./out/pricing.md', ['Pro is $10 a month']);
The package's main entry point does not load ws: the bridge runs in the daemon, which is a
separate process. @page-scanner/cli/daemon is the entry point that starts one in-process, and it is
the only one that pulls the WebSocket server in.
Limits
- No raster fallback. The extension can stitch screenshots when Chrome's debugger cannot attach,
but that path needs a user gesture, so a scan started from here fails loudly instead.
chrome.debugger cannot attach while DevTools is open on the same tab. Close it, or scan a
different tab.
- A page larger than 60,000 CSS pixels on a side is captured up to that limit, and the result says
so in
truncated. The extension's settings can raise the height's limit to 120,000 or
240,000, and a scan from here follows it.
- One pairing per machine, shared by every browser profile that connects to it.
- The helper needs Node 22 or later, for its WebSocket client. On Windows it is set up through the
registry, and has not yet been run on a Windows machine.
design reads each page's top document; a frame is not read. Its contrast is against the background an element
paints or inherits, so text over a layer another element paints is measured against the wrong
ground, and text over a gradient or a picture is listed for you to check by eye.
License
Apache-2.0. See LICENSE.