
Security News
/Company News
Securing the Financial Frontier: How Capital One Uses Socket for Open Source Security
Capital One is partnering with Socket to proactively secure its open source supply chain.
@page-scanner/cli
Advanced tools
Capture a whole web page from the command line, through the Page Scanner Chrome extension.
Capture a whole web page from the command line, as a PDF whose text is still text.
npx @page-scanner/cli install
npx @page-scanner/cli scan --url https://en.wikipedia.org/wiki/PDF --out ./pdf.pdf
The capture is done by the Page Scanner Chrome extension, in a Chrome you are already signed
in to. This package is the other half: it listens on loopback for that extension to connect, and
turns scan into a file on your disk. A page behind a login is a page you are logged into, and no
second browser is started.
It is also the engine behind @page-scanner/mcp, which is the same
operations as MCP tools for an agent.
--hide and no design.Two ways, in the order most people want them:
npx @page-scanner/cli --help # no install
npm install -g @page-scanner/cli # then just `page-scanner`
Once per machine:
page-scanner install
It sets up Page Scanner's helper for every Chromium browser it finds (Chrome, Chromium, Edge, Brave, Arc, Vivaldi, and Chrome's Beta, Dev and Canary), then says which. Open the Page Scanner settings in Chrome (the gear in the editor toolbar), go to Local Agents, press Connect, and allow Chrome's prompt. That is all: there is no port and no token to copy.
What it writes: a Native Messaging host manifest in each browser's NativeMessagingHosts
directory (on Windows, a registry value under HKCU), and the helper itself in
~/.page-scanner/native-host/. Chrome starts the helper for the Page Scanner extension and no
other; the helper reads the pairing in ~/.page-scanner/config.json, which install creates if
there is none, and connects the extension to the daemon below. page-scanner uninstall takes it
all out again and leaves the pairing.
install reads each profile's
list of extensions (and nothing else in it) for a copy of Page Scanner loaded from disk, and
allows it too. Chrome writes a newly loaded one down about 10 seconds after loading it, so run
install again if it was that quick. --extension-id <id> adds an id by hand.--browser-dir <dir> sets up only that user data directory, for a profile started with
--user-data-dir.install, because Chrome starts it without your shell's
PATH. status says when that Node is gone.--node <path> names another Node for the helper. Node reports its own path with symlinks
resolved, so a Homebrew Node is named by its versioned Cellar path, which the next
brew upgrade removes; --node /opt/homebrew/bin/node follows every upgrade instead.Chrome runs a separate copy of the extension in every profile, so your work profile and your personal profile are two different browsers here. Naming them in the settings is how you tell them apart later.
The older way, for a machine where the helper cannot be installed, and still working for every browser paired this way:
page-scanner pair
It prints a port and a token and saves them to ~/.page-scanner/config.json. In the settings'
Local Agents, set Connect through to A port and a token, paste both, name the browser,
and press Connect. pair waits up to two minutes and tells you when the browser arrives.
page-scanner pair --rotate issues a new token and invalidates the old one. Every browser paired
by hand has to be re-paired after that; the helper reads the new one by itself.
npx @page-scanner/cli asks npm for the newest version each time it runs, and installs it when it
is newer, so there is nothing to update. A global install stays on the version it installed, and
while it is there npx runs that copy too:
npm install -g @page-scanner/cli@latest
The next command that finds the old daemon still running shuts it down and starts its own. The
pairing in ~/.page-scanner/config.json carries over, and the extension reconnects by itself, so
nothing needs pairing again. The helper install copied does not need updating either: it only
carries frames, and it reaches whichever daemon is running.
If you also use the MCP server, keep the two on the same version. A command replaces a daemon of any other version, or of another build of the same version (which happens when the CLI is built from source), so two versions on one machine keep replacing each other's daemon, and the browser drops its connection each time.
page-scanner install [--extension-id <id>]... [--browser-dir <dir>]... [--node <path>] [--json]
page-scanner uninstall [--browser-dir <dir>]... [--json]
page-scanner pair [--port <n>] [--rotate] [--wait <s>=120] [--no-wait] [--json]
page-scanner status [--json]
page-scanner browsers [--json]
page-scanner tabs [--browser <id|label>] [--wait <s>=30] [--json]
page-scanner scan (--url <u>... | --urls <file|-> | --tab <id>) [--window <id>]
[--name <template>] [--markdown beside|only]
[--main-content] [--max-chars <n>]
[--links inline|references|text] [--pictures omit|alt]
[--redact]
[--hide <kinds>|all|none]
[--slices] [--slice-side <px>=1568] [--picture-files]
[--highlight <quote>]... [--highlight-from <file|->]
[--browser <id|label>]
[--format pdf|png|jpeg=pdf] [--page-size auto|a4|letter|phone=a4]
[--quality <0-1>] [--video frame|blank]
[--scheme auto|light|dark]
[--page-width window|a4|letter|phone] [--open-editor]
[--out <file|dir>] [--wait <s>=30] [--timeout <s>=120] [--json]
page-scanner diff <old> <new> [--section <heading>] [--out <file>] [--json]
page-scanner verify <file> [--record <file.integrity.json>] [--json]
page-scanner check <file.md> [<quote>...] [--from <file|->] [--loose] [--json]
page-scanner design (--url <u>... | --urls <file|-> | --tab <id> |
--crawl <u> [--max-pages <n>=10] [--depth <n>=2])
[--components] [--window <id>]
[--out <dir>] [--min-uses <n>=2] [--browser <id|label>]
[--wait <s>=30] [--timeout <s>=120] [--json]
page-scanner serve [--daemon] [--idle <min>]
page-scanner stop [--json]
page-scanner --version | --help
--browser takes a browserId or a label, matched without regard to case, and is only needed when
more than one browser is connected.
--out is a file when it ends in a 2 to 5 character extension and is not an existing directory,
and a directory otherwise, in which case the browser's suggested filename is used inside it.
Parent directories are created. A relative path is relative to your working directory, not the
daemon's, which is why the daemon hands the bytes back rather than writing them itself.
--page-width lays the page out at the width of a sheet before capturing it, the way a narrow
window would, so the PDF prints at 1:1 rather than being scaled down. window, the default, is
the width the browser has the page at. A 1280 px window on A4 is scaled to 55 %, which puts 16 px
body text at 6.5 pt; --page-width a4 puts it at 12. --page-width phone lays it out 390 px wide, as a phone shows it,
and --page-size phone cuts a PDF into phone screens, 390 × 844 px, with no margin.
--scheme picks which of a page's two themes to capture, for a page that has both. auto, the
default, is whichever the browser is showing; light and dark force one. It applies to the PDF
as well as the preview, which is the point: Chrome prints every page in the light scheme unless it
is told otherwise, so without this a dark page produced a white PDF.
--page-size decides what a PDF is laid onto. a4 and letter slice the capture across printable
sheets with a half-inch margin; auto is one page the exact size of the capture, which is not
printable and which Acrobat clamps past 200 inches.
--url more than once, or --urls <file> (one address to a line, # for a comment, - for
standard input), scans a list, one page at a time, into --out, which is then a directory. Each
path is printed as its file is written; a page that fails is reported on stderr and the rest carry
on, and the exit code is 1 if any failed. Up to 1,000 addresses a run.
--name names the file inside --out: {n} (the page's place in the list), {host}, {name}
(the browser's suggested name), {date}, {time} (when the run started) and {ext}, which is
added when left out. A / makes a subdirectory:
page-scanner scan --urls reading.txt --out ~/Captures/ --name '{date}/{n}-{host}'
--markdown beside also writes the page's text as a .md next to the file, read from the page
itself rather than from the PDF, and prints its path on a second line; --markdown only writes the
.md alone. With --json, the result carries the title, address, capture time, language and
headings as page. Pictures are left out of the text.
For an agent's context window, --main-content writes only the main content the extension found
(the whole page when it found none), --max-chars <n> keeps whole blocks up to that many
characters and reports in page.cut how long the whole was and how many blocks it left out,
--links references lists each address once at the end and --links text drops them, and
--pictures alt leaves a placeholder with each picture's alt text. Each heading in page.headings
has offset and chars, where its section starts in the text and how long it is. Dropping the
addresses is what shortens a link-heavy page most: Korean Wikipedia's "PDF" went from 85,657
characters to 37,134.
--redact replaces what Sensitive Text's rules find in the text before it leaves the browser:
every email address, phone number, card and IBAN, ID or tax number, IP address, key and token,
and the patterns you saved in the extension's Sensitive Text settings, each with a numbered
placeholder such as ⟦EMAIL 1⟧, the same value always the same placeholder. page.redacted counts
what was replaced by kind, never the values. It is for a signed-in page whose text goes on to a
cloud model; only the Markdown is redacted, so use it with --markdown only, since a PDF beside
it holds the page as it was. An extension older than this refuses the scan rather than send the
values.
--hide ads,consent,chat,overlays (or all, or none) hides ads, cookie and consent banners,
chat widgets and pop-ups before the capture and puts them back after; without it the extension's
own Clean up before capture settings decide. What was hidden is said on stderr and counted in
--json as hidden.
--slices also writes the capture as PNG slices for a vision model, top to bottom, into a folder
beside the file (page.pdf gives page.slices/slice-01.png and on). Each is at most
--slice-side pixels on a side, 1568 by default, the size past which Claude shrinks a picture,
and repeats the last 48 px of the one before, so a line on a seam is whole in one. A screenshot of
a whole long page is shrunk until its text cannot be read; a slice is not. --picture-files
writes the page's pictures and canvases (at least 120 by 40 px, up to 30 and 50) into
page.pictures/, cut from the capture where they were drawn. --json lists each file with the
band or box of the page it shows, in CSS px, as slices and pictures; slicesCut is set when a
page runs past 60 slices.
--highlight "<quote>", once a passage, or --highlight-from <file> (one a line, or a JSON
array; - reads standard input), marks passages quoted from the page where the page has them as
written, found the way check finds them, and a PDF lists each under a Quoted passages
bookmark. --json says which were found as highlighted; one not found is named on stderr and not
marked. It is for one page, not a list.
When the page declares how to cite it (citation tags, schema.org data, Dublin Core, OpenGraph),
--json has it as citation: the kind of work, and the title, authors, published date,
container (journal or site), volume, issue, pages, doi and the rest it gave, as it wrote
them and unchecked, with from, where they came from. The file itself carries the same. A page the
browser had machine-translated before the scan has translated ({ "by": "chrome" }, or
edge), and the file says it is a translation.
page-scanner diff <old> <new> says what changed between two captures, with no browser: two
Markdown files (or two PDFs with the .md beside them) a passage at a time, printed as diff -u
prints it, or two PNGs pixel by pixel, writing the newer one with the changed regions outlined. It
exits 0 when nothing changed and 1 when something did. --section "Pricing > Pro" compares only
the section under that heading, so a banner or a rail of latest posts around it does not count.
page-scanner verify <file> checks a file against the integrity record the extension wrote beside
it: the SHA-256 and size, and the signature when there is one. It exits 0 when both hold and 1 when
either does not.
page-scanner check <file.md> "<quote>"... says whether each quote or value an agent is about to
use occurs in a capture's Markdown as written, and where: the line and the section (the headings
above it, or front matter). The Markdown's own markup is undone first (escapes, emphasis, link
syntax, list and quote prefixes, table pipes) and line wrapping is ignored; a quote never runs
across two blocks, and a word or number has to stand apart (4.99 is not in $14.99). For a
missing quote it prints where the capture parts from it. A quote that matches only once curly
quotes, dashes and case are folded is reported apart, and --loose accepts it. Quotes come after
the file or from --from (one a line, or a JSON array; - reads standard input). No browser and
no model.
page-scanner design --url <address> reads a page's design as the browser draws
it and writes six files into --out, or design-<host> in the working directory. A design
system is spread over a site, so --url can be given more than once, or --urls <file|-> can
list up to 50 pages of it: they are read one after another and written as one design, and
audit.md adds the values only one of the pages uses. --crawl <address> finds the pages
instead: it follows the page's links to its own site, breadth first, reading one page of each
kind (one blog post, not all, but each section that has pages below it), up to --max-pages (10) and --depth links away (2). It
follows <a href> only, honors robots.txt and nofollow, and never opens an address that
acts, such as a log-out link, since the tabs are your own signed-in Chrome. A page that a
redirect took to another site, or to a page already read, is left out, and one the server
answered with an error (a 403 or a 404) is not read in. The six files: tokens.json
(W3C Design Tokens format), tokens.css, tailwind.preset.js, audit.md (values nearly equal,
probably meant as one), contrast.md (WCAG 2 contrast for every text color over its background)
and extract.json (the measurements, for an agent). The page is read in its light and its dark
scheme; an address opened for it is reloaded under each, a --tab you have open is not. Custom
properties keep their names, except a declared color nothing is drawn in, which is left out and
counted. Several names of one value, or of two colors nobody could tell apart, are one token under
the most used name, the others listed as its aliases, and a declared length counts only what its
name says it is for (--radius-* radii, --space-* spacing, --text-* font sizes); other values become tokens when at least --min-uses elements use
them, numbered by use. The folder's path goes to stdout and a summary to stderr. The
extension has to be one that says it can (an older one is refused with a hint to update).
--components finds the components as well: the buttons, fields and repeated boxes on the pages,
each with its variants (the instances that look alike) and what changes on hover and focus. Each
page is also scanned, and three more files are written: components.json, components.md, and
catalog.pdf, one section per component with each variant cut from the page as vector artwork
and its measured values beside it.
There is no scheduler here: cron, launchd and Task Scheduler run the command at a time. Start the
daemon at login with page-scanner serve --daemon --idle 0 so the extension is connected when the
schedule fires. The documentation has an example for each.
--wait 0 means do not wait at all. Chrome retires the extension's service worker after about
thirty seconds of silence, so the default wait exists to cover the reconnect that follows.
scan that is one line, the absolute path
of the file it wrote. For tabs and browsers it is a column-aligned table with no separator
row, so awk can read it.--json puts exactly one JSON document on stdout, for success and for failure alike, and
leaves stderr empty.{
"ok": true,
"path": "/Users/you/pdf.pdf",
"width": 1280,
"height": 4200,
"mode": "vector",
"selectableText": true,
"truncated": null,
"browserId": "b-9f2c41",
"fileName": "PDF - Wikipedia.pdf"
}
{ "ok": false, "code": 3, "error": "NO_BROWSER", "message": "No browser is connected." }
truncated is null rather than absent when the capture was whole, so a reader can see the
question was asked. When it is not null it carries what the page measured, what was captured, and
a sentence naming the gap.
| Code | Meaning |
|---|---|
| 0 | success |
| 1 | the browser was reached and the work failed |
| 2 | the arguments were wrong |
| 3 | no usable browser: none connected, several connected, or the one named is not |
| 4 | not paired |
| 5 | the daemon would not start |
diff uses the codes the way diff itself does: 0 when the two captures are the same, 1 when
something changed, and 2 for anything it cannot compare. check does the same: 0 when every quote is
there, 1 when one is not, 2 for a missing file or no quotes.
The extension is the WebSocket client: nothing outside Chrome can open a connection into a service worker, so something has to be listening before Chrome can dial in. That something is a background daemon, started on demand by the first command that needs it.
page-scanner scan ──HTTP /rpc──┐ ┌──< Chrome extension (paired by hand)
page-scanner-mcp ──HTTP /rpc──┼──> page-scanner daemon ──ws :45711
Claude Desktop ─HTTP /rpc───┘ (127.0.0.1) └──< helper ══stdio══ Chrome extension
The helper is what Chrome starts through Native Messaging. It dials 45711 the way the extension would, with the pairing's token, and waits for the daemon if none is running yet; so the daemon is still started by whatever wants a scan, and the browser arrives within a couple of seconds.
Two ports, deliberately. 45711 is the one the extension dials, and it refuses anything whose
Origin is not a chrome-extension://, which is exactly what a local CLI process is. The RPC port
is ephemeral, carries a bearer secret regenerated on every run, and is what the commands talk to.
page-scanner serve runs it in the foreground, which is the way to see why it will not start.
page-scanner stop stops it. It exits on its own after 15 minutes with no browser connected and
nothing calling (PAGE_SCANNER_IDLE_MINUTES, 0 for never), and never while a browser is
attached: the open socket is what keeps Chrome's service worker alive, and dropping it would make
the next command wait for a reconnect.
Everything lives in ~/.page-scanner, or in $PAGE_SCANNER_HOME if that is set.
| File | What it is |
|---|---|
config.json | the pairing: { port, token }, mode 0600 |
daemon.json | the running daemon: pid, RPC port, secret, version. Mode 0600 |
daemon.log | stdout and stderr of a daemon started in the background |
native-host/ | the helper Chrome starts, and the script that starts it |
Windows: those 0600 modes are a no-op. NTFS does not implement POSIX permissions, so both files are readable by anything running as you. That is usually fine on a personal machine. On a shared one, restrict the directory yourself:
icacls %USERPROFILE%\.page-scanner /inheritance:r /grant:r %USERNAME%:F
The token is the only thing between a local process and a scan of any tab you have open, because a loopback port is reachable by everything running as you. So:
127.0.0.1 only, never 0.0.0.0.chrome-extension:// Origin and the shared token in its
first frame. The Origin check alone is not enough: a local process can forge that header.fetch is blocked for an extension origin by
Local Network Access, so probing for a port would need a host permission the extension does not
have. The helper reads the port from the pairing file instead.install found, and any --extension-id you added. A store
extension of another id is never added by the search, whatever it is called. The helper holds no secret of its own; it reads the pairing
from a file only you can read, the same file pair writes.pair, on stdout, and never appears in any other output or in the log.chrome:// pages, the Web Store, and
the PDF viewer.import { scan, scanMany, listTabs, checkQuotesInFile } from '@page-scanner/cli';
const { tabs } = await listTabs();
const result = await scan({ tabId: tabs[0].tabId, out: './out/', format: 'pdf' });
console.log(result.path, result.selectableText);
// A list, one page at a time, with a result for each.
const batch = await scanMany(['https://example.com/a', 'https://example.com/b'], { out: './out/' });
// Whether an agent's quotes are in the capture, and where.
const { allFound, quotes } = checkQuotesInFile('./out/pricing.md', ['Pro is $10 a month']);
The package's main entry point does not load ws: the bridge runs in the daemon, which is a
separate process. @page-scanner/cli/daemon is the entry point that starts one in-process, and it is
the only one that pulls the WebSocket server in.
chrome.debugger cannot attach while DevTools is open on the same tab. Close it, or scan a
different tab.truncated. The extension's settings can raise the height's limit to 120,000 or
240,000, and a scan from here follows it.design reads each page's top document; a frame is not read. Its contrast is against the background an element
paints or inherits, so text over a layer another element paints is measured against the wrong
ground, and text over a gradient or a picture is listed for you to check by eye.Apache-2.0. See LICENSE.
FAQs
Capture a whole web page from the command line, through the Page Scanner Chrome extension.
The npm package @page-scanner/cli receives a total of 777 weekly downloads. As such, @page-scanner/cli popularity was classified as not popular.
We found that @page-scanner/cli demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Security News
/Company News
Capital One is partnering with Socket to proactively secure its open source supply chain.

Security News
Socket CTO Ahmad Nassri discusses how to keep AI agents from bypassing package blocks, limit credential access, and monitor their actions.

Security News
GPT-6 Astra tried to plant malicious code in simulated open source projects using fake GitHub accounts and deceptive PRs during an assigned CTF challenge.