New:Microsoft Teams Notifications Are Now Available in Socket.Learn more →
Get Started

@page-scanner/cli

Package Overview
Dependencies
Maintainers
1
Versions
10
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

@page-scanner/cli

Capture a whole web page from the command line, through the Page Scanner Chrome extension.

Source
npmnpm
Version
0.3.3
Version published
Weekly downloads
1.2K
4941.67%
Maintainers
1
Weekly downloads
 
Created
Source

page-scanner

Capture a whole web page from the command line, as a PDF whose text is still text.

npx @page-scanner/cli install
npx @page-scanner/cli scan --url https://en.wikipedia.org/wiki/PDF --out ./pdf.pdf

The capture is done by the Page Scanner Chrome extension, in a Chrome you are already signed in to. This package is the other half: it listens on loopback for that extension to connect, and turns scan into a file on your disk. A page behind a login is a page you are logged into, and no second browser is started.

It is also the engine behind @page-scanner/mcp, which is the same operations as MCP tools for an agent.

Requirements

  • Node.js 24 or newer.
  • Google Chrome, or another Chromium browser, with the Page Scanner extension 1.3.0 or newer installed. Earlier versions have no helper to connect through, no --hide and no design.

Install

Two ways, in the order most people want them:

npx @page-scanner/cli --help         # no install
npm install -g @page-scanner/cli     # then just `page-scanner`

Setting up

Once per machine:

page-scanner install

It sets up Page Scanner's helper for every Chromium browser it finds (Chrome, Chromium, Edge, Brave, Arc, Vivaldi, and Chrome's Beta, Dev and Canary), then says which. Open the Page Scanner settings in Chrome (the gear in the editor toolbar), go to Local Agents, press Connect, and allow Chrome's prompt. That is all: there is no port and no token to copy.

What it writes: a Native Messaging host manifest in each browser's NativeMessagingHosts directory (on Windows, a registry value under HKCU), and the helper itself in ~/.page-scanner/native-host/. Chrome starts the helper for the Page Scanner extension and no other; the helper reads the pairing in ~/.page-scanner/config.json, which install creates if there is none, and connects the extension to the daemon below. page-scanner uninstall takes it all out again and leaves the pairing.

  • An unpacked build loaded in developer mode is found by itself: install reads each profile's list of extensions (and nothing else in it) for a copy of Page Scanner loaded from disk, and allows it too. Chrome writes a newly loaded one down about 10 seconds after loading it, so run install again if it was that quick. --extension-id <id> adds an id by hand.
  • --browser-dir <dir> sets up only that user data directory, for a profile started with --user-data-dir.
  • Run it again after installing another browser, or after removing the Node it names: the helper is started with the Node that ran install, because Chrome starts it without your shell's PATH. status says when that Node is gone.
  • --node <path> names another Node for the helper. Node reports its own path with symlinks resolved, so a Homebrew Node is named by its versioned Cellar path, which the next brew upgrade removes; --node /opt/homebrew/bin/node follows every upgrade instead.

Chrome runs a separate copy of the extension in every profile, so your work profile and your personal profile are two different browsers here. Naming them in the settings is how you tell them apart later.

Pairing with a port and a token

The older way, for a machine where the helper cannot be installed, and still working for every browser paired this way:

page-scanner pair

It prints a port and a token and saves them to ~/.page-scanner/config.json. In the settings' Local Agents, set Connect through to A port and a token, paste both, name the browser, and press Connect. pair waits up to two minutes and tells you when the browser arrives.

page-scanner pair --rotate issues a new token and invalidates the old one. Every browser paired by hand has to be re-paired after that; the helper reads the new one by itself.

Updating

npx @page-scanner/cli asks npm for the newest version each time it runs, and installs it when it is newer, so there is nothing to update. A global install stays on the version it installed, and while it is there npx runs that copy too:

npm install -g @page-scanner/cli@latest

The next command that finds the old daemon still running shuts it down and starts its own. The pairing in ~/.page-scanner/config.json carries over, and the extension reconnects by itself, so nothing needs pairing again. The helper install copied does not need updating either: it only carries frames, and it reaches whichever daemon is running.

If you also use the MCP server, keep the two on the same version. A command replaces a daemon of any other version, or of another build of the same version (which happens when the CLI is built from source), so two versions on one machine keep replacing each other's daemon, and the browser drops its connection each time.

Commands

page-scanner install  [--extension-id <id>]... [--browser-dir <dir>]... [--node <path>] [--json]
page-scanner uninstall [--browser-dir <dir>]... [--json]
page-scanner pair     [--port <n>] [--rotate] [--wait <s>=120] [--no-wait] [--json]
page-scanner status   [--json]
page-scanner browsers [--json]
page-scanner tabs     [--browser <id|label>] [--wait <s>=30] [--json]
page-scanner scan     (--url <u>... | --urls <file|-> | --tab <id>) [--window <id>]
                      [--name <template>] [--markdown beside|only]
                      [--main-content] [--max-chars <n>]
                      [--links inline|references|text] [--pictures omit|alt]
                      [--redact]
                      [--hide <kinds>|all|none]
                      [--slices] [--slice-side <px>=1568] [--picture-files]
                      [--highlight <quote>]... [--highlight-from <file|->]
                      [--browser <id|label>]
                      [--format pdf|png|jpeg=pdf] [--page-size auto|a4|letter|phone=a4]
                      [--quality <0-1>] [--video frame|blank]
                      [--scheme auto|light|dark]
                      [--page-width window|a4|letter|phone] [--open-editor]
                      [--out <file|dir>] [--wait <s>=30] [--timeout <s>=120] [--json]
page-scanner diff     <old> <new> [--section <heading>] [--out <file>] [--json]
page-scanner verify   <file> [--record <file.integrity.json>] [--json]
page-scanner check    <file.md> [<quote>...] [--from <file|->] [--loose] [--json]
page-scanner design   (--url <u>... | --urls <file|-> | --tab <id> |
                       --crawl <u> [--max-pages <n>=10] [--depth <n>=2])
                      [--components] [--window <id>]
                      [--out <dir>] [--min-uses <n>=2] [--browser <id|label>]
                      [--wait <s>=30] [--timeout <s>=120] [--json]
page-scanner serve    [--daemon] [--idle <min>]
page-scanner stop     [--json]
page-scanner --version | --help

--browser takes a browserId or a label, matched without regard to case, and is only needed when more than one browser is connected.

--out is a file when it ends in a 2 to 5 character extension and is not an existing directory, and a directory otherwise, in which case the browser's suggested filename is used inside it. Parent directories are created. A relative path is relative to your working directory, not the daemon's, which is why the daemon hands the bytes back rather than writing them itself.

--page-width lays the page out at the width of a sheet before capturing it, the way a narrow window would, so the PDF prints at 1:1 rather than being scaled down. window, the default, is the width the browser has the page at. A 1280 px window on A4 is scaled to 55 %, which puts 16 px body text at 6.5 pt; --page-width a4 puts it at 12. --page-width phone lays it out 390 px wide, as a phone shows it, and --page-size phone cuts a PDF into phone screens, 390 × 844 px, with no margin.

--scheme picks which of a page's two themes to capture, for a page that has both. auto, the default, is whichever the browser is showing; light and dark force one. It applies to the PDF as well as the preview, which is the point: Chrome prints every page in the light scheme unless it is told otherwise, so without this a dark page produced a white PDF.

--page-size decides what a PDF is laid onto. a4 and letter slice the capture across printable sheets with a half-inch margin; auto is one page the exact size of the capture, which is not printable and which Acrobat clamps past 200 inches.

--url more than once, or --urls <file> (one address to a line, # for a comment, - for standard input), scans a list, one page at a time, into --out, which is then a directory. Each path is printed as its file is written; a page that fails is reported on stderr and the rest carry on, and the exit code is 1 if any failed. Up to 1,000 addresses a run.

--name names the file inside --out: {n} (the page's place in the list), {host}, {name} (the browser's suggested name), {date}, {time} (when the run started) and {ext}, which is added when left out. A / makes a subdirectory:

page-scanner scan --urls reading.txt --out ~/Captures/ --name '{date}/{n}-{host}'

--markdown beside also writes the page's text as a .md next to the file, read from the page itself rather than from the PDF, and prints its path on a second line; --markdown only writes the .md alone. With --json, the result carries the title, address, capture time, language and headings as page. Pictures are left out of the text.

For an agent's context window, --main-content writes only the main content the extension found (the whole page when it found none), --max-chars <n> keeps whole blocks up to that many characters and reports in page.cut how long the whole was and how many blocks it left out, --links references lists each address once at the end and --links text drops them, and --pictures alt leaves a placeholder with each picture's alt text. Each heading in page.headings has offset and chars, where its section starts in the text and how long it is. Dropping the addresses is what shortens a link-heavy page most: Korean Wikipedia's "PDF" went from 85,657 characters to 37,134.

--redact replaces what Sensitive Text's rules find in the text before it leaves the browser: every email address, phone number, card and IBAN, ID or tax number, IP address, key and token, and the patterns you saved in the extension's Sensitive Text settings, each with a numbered placeholder such as ⟦EMAIL 1⟧, the same value always the same placeholder. page.redacted counts what was replaced by kind, never the values. It is for a signed-in page whose text goes on to a cloud model; only the Markdown is redacted, so use it with --markdown only, since a PDF beside it holds the page as it was. An extension older than this refuses the scan rather than send the values.

--hide ads,consent,chat,overlays (or all, or none) hides ads, cookie and consent banners, chat widgets and pop-ups before the capture and puts them back after; without it the extension's own Clean up before capture settings decide. What was hidden is said on stderr and counted in --json as hidden.

--slices also writes the capture as PNG slices for a vision model, top to bottom, into a folder beside the file (page.pdf gives page.slices/slice-01.png and on). Each is at most --slice-side pixels on a side, 1568 by default, the size past which Claude shrinks a picture, and repeats the last 48 px of the one before, so a line on a seam is whole in one. A screenshot of a whole long page is shrunk until its text cannot be read; a slice is not. --picture-files writes the page's pictures and canvases (at least 120 by 40 px, up to 30 and 50) into page.pictures/, cut from the capture where they were drawn. --json lists each file with the band or box of the page it shows, in CSS px, as slices and pictures; slicesCut is set when a page runs past 60 slices.

--highlight "<quote>", once a passage, or --highlight-from <file> (one a line, or a JSON array; - reads standard input), marks passages quoted from the page where the page has them as written, found the way check finds them, and a PDF lists each under a Quoted passages bookmark. --json says which were found as highlighted; one not found is named on stderr and not marked. It is for one page, not a list.

When the page declares how to cite it (citation tags, schema.org data, Dublin Core, OpenGraph), --json has it as citation: the kind of work, and the title, authors, published date, container (journal or site), volume, issue, pages, doi and the rest it gave, as it wrote them and unchecked, with from, where they came from. The file itself carries the same. A page the browser had machine-translated before the scan has translated ({ "by": "chrome" }, or edge), and the file says it is a translation.

page-scanner diff <old> <new> says what changed between two captures, with no browser: two Markdown files (or two PDFs with the .md beside them) a passage at a time, printed as diff -u prints it, or two PNGs pixel by pixel, writing the newer one with the changed regions outlined. It exits 0 when nothing changed and 1 when something did. --section "Pricing > Pro" compares only the section under that heading, so a banner or a rail of latest posts around it does not count.

page-scanner verify <file> checks a file against the integrity record the extension wrote beside it: the SHA-256 and size, and the signature when there is one. It exits 0 when both hold and 1 when either does not.

page-scanner check <file.md> "<quote>"... says whether each quote or value an agent is about to use occurs in a capture's Markdown as written, and where: the line and the section (the headings above it, or front matter). The Markdown's own markup is undone first (escapes, emphasis, link syntax, list and quote prefixes, table pipes) and line wrapping is ignored; a quote never runs across two blocks, and a word or number has to stand apart (4.99 is not in $14.99). For a missing quote it prints where the capture parts from it. A quote that matches only once curly quotes, dashes and case are folded is reported apart, and --loose accepts it. Quotes come after the file or from --from (one a line, or a JSON array; - reads standard input). No browser and no model.

page-scanner design --url <address> reads a page's design as the browser draws it and writes six files into --out, or design-<host> in the working directory. A design system is spread over a site, so --url can be given more than once, or --urls <file|-> can list up to 50 pages of it: they are read one after another and written as one design, and audit.md adds the values only one of the pages uses. --crawl <address> finds the pages instead: it follows the page's links to its own site, breadth first, reading one page of each kind (one blog post, not all, but each section that has pages below it), up to --max-pages (10) and --depth links away (2). It follows <a href> only, honors robots.txt and nofollow, and never opens an address that acts, such as a log-out link, since the tabs are your own signed-in Chrome. A page that a redirect took to another site, or to a page already read, is left out, and one the server answered with an error (a 403 or a 404) is not read in. The six files: tokens.json (W3C Design Tokens format), tokens.css, tailwind.preset.js, audit.md (values nearly equal, probably meant as one), contrast.md (WCAG 2 contrast for every text color over its background) and extract.json (the measurements, for an agent). The page is read in its light and its dark scheme; an address opened for it is reloaded under each, a --tab you have open is not. Custom properties keep their names, except a declared color nothing is drawn in, which is left out and counted. Several names of one value, or of two colors nobody could tell apart, are one token under the most used name, the others listed as its aliases, and a declared length counts only what its name says it is for (--radius-* radii, --space-* spacing, --text-* font sizes); other values become tokens when at least --min-uses elements use them, numbered by use. The folder's path goes to stdout and a summary to stderr. The extension has to be one that says it can (an older one is refused with a hint to update).

--components finds the components as well: the buttons, fields and repeated boxes on the pages, each with its variants (the instances that look alike) and what changes on hover and focus. Each page is also scanned, and three more files are written: components.json, components.md, and catalog.pdf, one section per component with each variant cut from the page as vector artwork and its measured values beside it.

There is no scheduler here: cron, launchd and Task Scheduler run the command at a time. Start the daemon at login with page-scanner serve --daemon --idle 0 so the extension is connected when the schedule fires. The documentation has an example for each.

--wait 0 means do not wait at all. Chrome retires the extension's service worker after about thirty seconds of silence, so the default wait exists to cover the reconnect that follows.

What a command prints

  • stdout carries the answer and nothing else. For scan that is one line, the absolute path of the file it wrote. For tabs and browsers it is a column-aligned table with no separator row, so awk can read it.
  • stderr carries anything addressed to a person: progress, warnings, the reason something failed.
  • --json puts exactly one JSON document on stdout, for success and for failure alike, and leaves stderr empty.
{
  "ok": true,
  "path": "/Users/you/pdf.pdf",
  "width": 1280,
  "height": 4200,
  "mode": "vector",
  "selectableText": true,
  "truncated": null,
  "browserId": "b-9f2c41",
  "fileName": "PDF - Wikipedia.pdf"
}
{ "ok": false, "code": 3, "error": "NO_BROWSER", "message": "No browser is connected." }

truncated is null rather than absent when the capture was whole, so a reader can see the question was asked. When it is not null it carries what the page measured, what was captured, and a sentence naming the gap.

Exit codes

CodeMeaning
0success
1the browser was reached and the work failed
2the arguments were wrong
3no usable browser: none connected, several connected, or the one named is not
4not paired
5the daemon would not start

diff uses the codes the way diff itself does: 0 when the two captures are the same, 1 when something changed, and 2 for anything it cannot compare. check does the same: 0 when every quote is there, 1 when one is not, 2 for a missing file or no quotes.

The daemon

The extension is the WebSocket client: nothing outside Chrome can open a connection into a service worker, so something has to be listening before Chrome can dial in. That something is a background daemon, started on demand by the first command that needs it.

page-scanner scan ──HTTP /rpc──┐                           ┌──< Chrome extension (paired by hand)
page-scanner-mcp  ──HTTP /rpc──┼──> page-scanner daemon ──ws :45711
Claude Desktop    ─HTTP /rpc───┘    (127.0.0.1)            └──< helper ══stdio══ Chrome extension

The helper is what Chrome starts through Native Messaging. It dials 45711 the way the extension would, with the pairing's token, and waits for the daemon if none is running yet; so the daemon is still started by whatever wants a scan, and the browser arrives within a couple of seconds.

Two ports, deliberately. 45711 is the one the extension dials, and it refuses anything whose Origin is not a chrome-extension://, which is exactly what a local CLI process is. The RPC port is ephemeral, carries a bearer secret regenerated on every run, and is what the commands talk to.

page-scanner serve runs it in the foreground, which is the way to see why it will not start. page-scanner stop stops it. It exits on its own after 15 minutes with no browser connected and nothing calling (PAGE_SCANNER_IDLE_MINUTES, 0 for never), and never while a browser is attached: the open socket is what keeps Chrome's service worker alive, and dropping it would make the next command wait for a reconnect.

Files

Everything lives in ~/.page-scanner, or in $PAGE_SCANNER_HOME if that is set.

FileWhat it is
config.jsonthe pairing: { port, token }, mode 0600
daemon.jsonthe running daemon: pid, RPC port, secret, version. Mode 0600
daemon.logstdout and stderr of a daemon started in the background
native-host/the helper Chrome starts, and the script that starts it

Windows: those 0600 modes are a no-op. NTFS does not implement POSIX permissions, so both files are readable by anything running as you. That is usually fine on a personal machine. On a shared one, restrict the directory yourself:

icacls %USERPROFILE%\.page-scanner /inheritance:r /grant:r %USERNAME%:F

Security

The token is the only thing between a local process and a scan of any tab you have open, because a loopback port is reachable by everything running as you. So:

  • The bridge binds 127.0.0.1 only, never 0.0.0.0.
  • Every connection must present a chrome-extension:// Origin and the shared token in its first frame. The Origin check alone is not enough: a local process can forge that header.
  • The extension does no port discovery. Loopback fetch is blocked for an extension origin by Local Network Access, so probing for a port would need a host permission the extension does not have. The helper reads the port from the pairing file instead.
  • Chrome starts the helper only for the extension ids its manifest names: the two store builds' (the Chrome Web Store's and Edge Add-ons'), the unpacked copies of Page Scanner install found, and any --extension-id you added. A store extension of another id is never added by the search, whatever it is called. The helper holds no secret of its own; it reads the pairing from a file only you can read, the same file pair writes.
  • The token is printed by pair, on stdout, and never appears in any other output or in the log.
  • The extension refuses tabs Chrome will not let it touch: chrome:// pages, the Web Store, and the PDF viewer.

Using it from Node

import { scan, scanMany, listTabs, checkQuotesInFile } from '@page-scanner/cli';

const { tabs } = await listTabs();
const result = await scan({ tabId: tabs[0].tabId, out: './out/', format: 'pdf' });
console.log(result.path, result.selectableText);

// A list, one page at a time, with a result for each.
const batch = await scanMany(['https://example.com/a', 'https://example.com/b'], { out: './out/' });

// Whether an agent's quotes are in the capture, and where.
const { allFound, quotes } = checkQuotesInFile('./out/pricing.md', ['Pro is $10 a month']);

The package's main entry point does not load ws: the bridge runs in the daemon, which is a separate process. @page-scanner/cli/daemon is the entry point that starts one in-process, and it is the only one that pulls the WebSocket server in.

Limits

  • No raster fallback. The extension can stitch screenshots when Chrome's debugger cannot attach, but that path needs a user gesture, so a scan started from here fails loudly instead.
  • chrome.debugger cannot attach while DevTools is open on the same tab. Close it, or scan a different tab.
  • A page larger than 60,000 CSS pixels on a side is captured up to that limit, and the result says so in truncated. The extension's settings can raise the height's limit to 120,000 or 240,000, and a scan from here follows it.
  • One pairing per machine, shared by every browser profile that connects to it.
  • The helper needs Node 22 or later, for its WebSocket client. On Windows it is set up through the registry, and has not yet been run on a Windows machine.
  • design reads each page's top document; a frame is not read. Its contrast is against the background an element paints or inherits, so text over a layer another element paints is measured against the wrong ground, and text over a gradient or a picture is listed for you to check by eye.

License

Apache-2.0. See LICENSE.

Keywords

chrome

FAQs

Package last updated on 28 Sep 2026

Related posts