New:Microsoft Teams Notifications Are Now Available in Socket.Learn more →
Get Started

@page-scanner/cli

Package Overview
Dependencies
Maintainers
1
Versions
12
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

@page-scanner/cli

Capture a whole web page from the command line, through the Page Scanner Chrome extension.

Source
npmnpm
Version
0.1.0
Version published
Weekly downloads
314
-67.36%
Maintainers
1
Weekly downloads
 
Created
Source

page-scanner

Capture a whole web page from the command line, as a PDF whose text is still text.

npx @page-scanner/cli pair
npx @page-scanner/cli scan --url https://en.wikipedia.org/wiki/PDF --out ./pdf.pdf

The capture is done by the Page Scanner Chrome extension, in a Chrome you are already signed in to. This package is the other half: it listens on loopback for that extension to connect, and turns scan into a file on your disk. A page behind a login is a page you are logged into, and no second browser is started.

It is also the engine behind @page-scanner/mcp, which is the same four operations as MCP tools for an agent.

Requirements

  • Node.js 24 or newer.
  • Google Chrome, or another Chromium browser, with the Page Scanner extension installed.

Install

Three ways, in the order most people want them:

npx @page-scanner/cli --help         # no install
npm install -g @page-scanner/cli     # then just `page-scanner`

Or a standalone binary from the releases page, which needs no Node at all. The binaries are not code-signed. On macOS, Gatekeeper will refuse the first run until you clear the quarantine flag:

xattr -d com.apple.quarantine ./page-scanner

On Windows, SmartScreen shows "Windows protected your PC"; the binary runs from More info then Run anyway. Verify what you downloaded against SHA256SUMS on the same release.

Pairing

Once per machine:

page-scanner pair

It prints a port and a token and saves them to ~/.page-scanner/config.json. Open the Page Scanner settings page in Chrome (the gear in the editor toolbar), paste both into Local agents, name the browser, and press Connect. pair waits up to two minutes and tells you when the browser arrives.

Chrome runs a separate copy of the extension in every profile, so your work profile and your personal profile are two different browsers here. Naming them is how you tell them apart later.

page-scanner pair --rotate issues a new token and invalidates the old one. Every browser has to be re-paired after that.

Commands

page-scanner pair     [--port <n>] [--rotate] [--wait <s>=120] [--no-wait] [--json]
page-scanner status   [--json]
page-scanner browsers [--json]
page-scanner tabs     [--browser <id|label>] [--wait <s>=30] [--json]
page-scanner scan     (--url <u> | --tab <id>) [--window <id>] [--browser <id|label>]
                      [--format pdf|png|jpeg=pdf] [--page-size auto|a4|letter=a4]
                      [--quality <0-1>] [--video frame|blank] [--open-editor]
                      [--out <file|dir>] [--wait <s>=30] [--timeout <s>=120] [--json]
page-scanner serve    [--daemon] [--idle <min>]
page-scanner stop
page-scanner --version | --help

--browser takes a browserId or a label, matched without regard to case, and is only needed when more than one browser is connected.

--out is a file when it ends in a 2 to 5 character extension and is not an existing directory, and a directory otherwise, in which case the browser's suggested filename is used inside it. Parent directories are created. A relative path is relative to your working directory, not the daemon's, which is why the daemon hands the bytes back rather than writing them itself.

--page-size decides what a PDF is laid onto. a4 and letter slice the capture across printable sheets with a half-inch margin; auto is one page the exact size of the capture, which is not printable and which Acrobat clamps past 200 inches.

--wait 0 means do not wait at all. Chrome retires the extension's service worker after about thirty seconds of silence, so the default wait exists to cover the reconnect that follows.

What a command prints

  • stdout carries the answer and nothing else. For scan that is one line, the absolute path of the file it wrote. For tabs and browsers it is a column-aligned table with no separator row, so awk can read it.
  • stderr carries anything addressed to a person: progress, warnings, the reason something failed.
  • --json puts exactly one JSON document on stdout, for success and for failure alike, and leaves stderr empty.
{
  "ok": true,
  "path": "/Users/you/pdf.pdf",
  "width": 1280,
  "height": 4200,
  "mode": "vector",
  "selectableText": true,
  "truncated": null,
  "browserId": "b-9f2c41",
  "fileName": "PDF - Wikipedia.pdf"
}
{ "ok": false, "code": 3, "error": "NO_BROWSER", "message": "No browser is connected." }

truncated is null rather than absent when the capture was whole, so a reader can see the question was asked. When it is not null it carries what the page measured, what was captured, and a sentence naming the gap.

Exit codes

CodeMeaning
0success
1the browser was reached and the work failed
2the arguments were wrong
3no usable browser: none connected, several connected, or the one named is not
4not paired
5the daemon would not start

The daemon

The extension is the WebSocket client: nothing outside Chrome can open a connection into a service worker, so something has to be listening before Chrome can dial in. That something is a background daemon, started on demand by the first command that needs it.

page-scanner scan ──HTTP /rpc──┐
page-scanner-mcp  ──HTTP /rpc──┼──> page-scanner daemon ──ws :45711──< Chrome extension(s)
Claude Desktop    ─HTTP /rpc───┘    (127.0.0.1, ephemeral RPC port)

Two ports, deliberately. 45711 is the one the extension dials, and it refuses anything whose Origin is not a chrome-extension://, which is exactly what a local CLI process is. The RPC port is ephemeral, carries a bearer secret regenerated on every run, and is what the commands talk to.

page-scanner serve runs it in the foreground, which is the way to see why it will not start. page-scanner stop stops it. It exits on its own after 15 minutes with no browser connected and nothing calling (PAGE_SCANNER_IDLE_MINUTES, 0 for never), and never while a browser is attached: the open socket is what keeps Chrome's service worker alive, and dropping it would make the next command wait for a reconnect.

Files

Everything lives in ~/.page-scanner, or in $PAGE_SCANNER_HOME if that is set.

FileWhat it is
config.jsonthe pairing: { port, token }, mode 0600
daemon.jsonthe running daemon: pid, RPC port, secret, version. Mode 0600
daemon.logstdout and stderr of a daemon started in the background

Windows: those 0600 modes are a no-op. NTFS does not implement POSIX permissions, so both files are readable by anything running as you. That is usually fine on a personal machine. On a shared one, restrict the directory yourself:

icacls %USERPROFILE%\.page-scanner /inheritance:r /grant:r %USERNAME%:F

Security

The token is the only thing between a local process and a scan of any tab you have open, because a loopback port is reachable by everything running as you. So:

  • The bridge binds 127.0.0.1 only, never 0.0.0.0.
  • Every connection must present a chrome-extension:// Origin and the shared token in its first frame. The Origin check alone is not enough: a local process can forge that header.
  • There is no port discovery. Loopback fetch is blocked for an extension origin by Local Network Access, so probing for a port would need a host permission the extension does not have.
  • The token is printed by pair, on stdout, and never appears in any other output or in the log.
  • The extension refuses tabs Chrome will not let it touch: chrome:// pages, the Web Store, and the PDF viewer.

Using it from Node

import { scan, listTabs } from '@page-scanner/cli';

const { tabs } = await listTabs();
const result = await scan({ tabId: tabs[0].tabId, out: './out/', format: 'pdf' });
console.log(result.path, result.selectableText);

The package's main entry point does not load ws: the bridge runs in the daemon, which is a separate process. @page-scanner/cli/daemon is the entry point that starts one in-process, and it is the only one that pulls the WebSocket server in.

Limits

  • No raster fallback. The extension can stitch screenshots when Chrome's debugger cannot attach, but that path needs a user gesture, so a scan started from here fails loudly instead.
  • chrome.debugger cannot attach while DevTools is open on the same tab. Close it, or scan a different tab.
  • A page larger than 60,000 CSS pixels on a side is captured up to that limit, and the result says so in truncated.
  • One pairing per machine, shared by every browser profile that connects to it.

Licence

Apache-2.0. See LICENSE.

Keywords

chrome

FAQs

Package last updated on 12 Sep 2026

Related posts