
Security News
Happy Birthday, Shai-Hulud
It has been one year since Shai-Hulud made its first appearance on npm.
scrapewright
Advanced tools
Give it a store URL — it detects the platform and, for custom sites, synthesizes a reusable scraper once via an LLM, then replays it for free.
Give it a URL. It writes the scraper.
Most e-commerce catalog scraping splits into two worlds: sites on a known platform (Shopify, WooCommerce) that expose a clean JSON feed, and everything else — bespoke HTML where you hand-write a parser per site and re-write it every time the markup shifts. scrapewright collapses both into one call:
The LLM is a compiler, not a runtime. It runs once per site to produce a recipe of CSS selectors; every page after that is parsed by plain BeautifulSoup at zero marginal cost. That is the whole cost-control story — no per-page model calls, no token bill that scales with your crawl.
┌─────────────┐
store URL ───▶ │ detect │
└──────┬──────┘
┌──────────────────┼──────────────────┐
▼ ▼ ▼
shopify woocommerce generic HTML
products.json wc/store/products (page mode)
│ │ │
│ deterministic │ ▼
│ (free) │ cached recipe? ──yes──▶ replay (free)
└────────┬─────────┘ │ no
▼ ▼
Product{} ◀───── selectors ── JSON-LD? ──yes──▶ Product{} (free)
▲ │ no
│ ▼
└──────── replay ◀── LLM synthesizes recipe ONCE ──▶ cache
Everything normalizes to one Product shape, so downstream code never knows or
cares which path a record came from.
pip install scrapewright # deterministic paths (Shopify, Woo, JSON-LD)
pip install "scrapewright[llm]" # + LLM recipe synthesis for custom HTML
pip install "scrapewright[llm,js,excel,mcp]" # + JS rendering, XLSX, MCP server
playwright install chromium # only needed for --js
from scrapewright import Scrapewright
sw = Scrapewright()
# Catalog mode — a whole Shopify/WooCommerce store, deterministically
for product in sw.scrape_catalog("https://shop.example.com", max_items=200):
print(product.brand, product.title, product.price, product.currency)
# Page mode — one custom-HTML product page.
# First call: tries JSON-LD (free); if absent, the LLM writes a recipe once.
# Every later call on that domain: replayed from the cached recipe, no LLM.
item = sw.scrape_page("https://boutique.example.com/products/wool-coat")
print(item.model_dump(exclude={"raw"}))
# Crawl mode — walk a WHOLE custom store from one listing/category URL.
# The frontier discovers product pages (deterministic, free); the first page
# pays the single synthesis cost, every other page replays the recipe.
for product in sw.crawl("https://boutique.example.com/collection", max_items=100):
print(product.title, product.price)
scrapewright detect https://shop.example.com # platform + strategy
scrapewright run https://shop.example.com --max 50 # scrape a catalog → JSONL
scrapewright crawl https://boutique.example.com/collection -o products.xlsx
scrapewright run https://shop.example.com -o products.csv # Excel-ready CSV
scrapewright add https://boutique.example.com/products/coat # learn a site
scrapewright run https://boutique.example.com/products/coat --no-llm
scrapewright list # cached recipe domains
-o writes .csv (Excel-ready, UTF-8 BOM), .xlsx (pip install scrapewright[excel]),
or .jsonl; without it, products stream to stdout as JSONL.
detect answers the routing question before a job starts:
$ scrapewright detect https://some-store.com
https://some-store.com
platform: bigcommerce
catalog: -
strategy: crawl
note: BigCommerce (Stencil) markup
Twelve platforms are recognized: Shopify and WooCommerce publish a free
JSON catalog, so those route to catalog — deterministic, no LLM, no browser.
Magento, BigCommerce, Salesforce Commerce Cloud, Squarespace, Wix, Webflow,
PrestaShop, Shopware, Ecwid and OpenCart are recognized by fingerprint and
route to crawl, where the recipe path handles them like any custom site — the
point of naming them is knowing what you face, not writing twelve parsers.
Wix and Ecwid render client-side, so detection says crawl+js up front.
A site behind an anti-bot wall reports strategy: blocked with the HTTP status,
rather than pretending it found nothing.
Products are just the built-in default. Declare the fields you want and the same compile-once/replay-free loop works on any structured page — job posts, listings, registry records:
scrapewright run https://jobs.example.com/p/123 -f title -f company -f salary:number -f tags:list --schema-name job
from scrapewright import Scrapewright, Schema
job = Schema.from_names(["title", "company", "salary:number", "tags:list"], name="job")
record = Scrapewright().extract("https://jobs.example.com/p/123", job)
print(record.data) # {'title': ..., 'company': ..., 'salary': ..., 'tags': [...]}
Field kinds are text (default), number, url, and list. Recipes are cached
per site and per schema, so one domain can be compiled against several field
sets without them overwriting each other.
scrapewright speaks MCP, so an agent can call it as a tool instead of reading raw HTML itself. Two ways in.
Hosted, nothing to install. Point the client at the service with a key from scrapewright.app (1,000 free rows a month):
{
"mcpServers": {
"scrapewright": {
"url": "https://scrapewright.app/mcp",
"headers": { "X-API-Key": "sw_..." }
}
}
}
Tools: detect_site, extract_page, crawl_site, crawl_status, account. Paid
from the same credit balance as the REST API; no model key of your own is needed.
Local, your own model key. Run the server on your machine:
pip install "scrapewright[mcp,llm]"
scrapewright mcp
Point any MCP client at that command and the agent gains five tools: detect_site,
scrape_catalog, extract_page, crawl_site, and list_learned_sites.
Drop this into your client's config — Claude Desktop, Cursor, or anything else that speaks MCP:
{
"mcpServers": {
"scrapewright": {
"command": "uvx",
"args": ["--from", "scrapewright[mcp,llm]", "scrapewright", "mcp"],
"env": { "ANTHROPIC_API_KEY": "sk-ant-..." }
}
}
}
The key is only needed for sites on no known platform, where a recipe has to be written once. Shopify and WooCommerce stores work without it.
The economics are the point. An agent that reads pages itself pays model tokens per page, forever. These tools pay once per site — an agent crawling 500 pages spends one synthesis, not five hundred, and platform stores (Shopify, WooCommerce) cost nothing at all.
The same core behind an HTTP API, with keys, quotas, metering and background jobs:
pip install "scrapewright[service,llm]"
scrapewright keys create --label alice --plan free
scrapewright serve --port 8000
curl -X POST localhost:8000/v1/extract -H "X-API-Key: sw_..." -H "Content-Type: application/json" -d '{"url": "https://shop.example.com/products/coat"}'
| Endpoint | Purpose |
|---|---|
POST /v1/detect | platform + strategy (cheap) |
POST /v1/extract | one page -> structured record |
POST /v1/crawl | a whole site -> job id (crawls outlive a request) |
GET /v1/jobs/{id} | poll a crawl |
GET /v1/usage | what this key has consumed, against its plan |
One action costs real money: compiling a new site, a single LLM pass over a page, measured at $0.02 on a small product page and $0.15 on a heavy rendered one. Everything after that is BeautifulSoup — the ten-thousandth record from a compiled site is free to serve. So credits are priced off that one action, and everything else is denominated relative to it:
| Action | Credits |
|---|---|
| 1 record delivered | 1 |
| 1 browser render | 5 |
| 1 new site compiled | 300 |
page fetches, detect | free |
$ scrapewright plans
pack credits price $/credit margin
starter 10,000 $10 0.00100 80.0%
growth 50,000 $40 0.00080 75.0%
scale 250,000 $150 0.00060 66.7%
Free: 1,000 credits a month, resetting.
Margin is measured on compiling a site, because that is the only step that costs anything; a test fails if a price edit drops any pack below 60%. A free account can cost us at most $0.20 a month, even if every free credit goes to the most expensive action there is.
Credits are a ledger, not a counter — every grant and every charge is a row,
so a disputed bill can be reconstructed line by line, and a replayed payment
webhook cannot double-credit (grants take an idempotency key). Running out
returns 402 with the balance and what to do about it; a crawl is capped by the
credits on hand, so a job stops at what the caller can pay for instead of
overdrawing.
scrapewright credits grant <key_id> --pack starter --idempotency <payment_id>
scrapewright credits balance <key_id>
Stripe is wired in and turned on by environment, not by a code change:
pip install "scrapewright[service,stripe]"
export STRIPE_SECRET_KEY=sk_test_... # absent -> nothing is for sale
export STRIPE_WEBHOOK_SECRET=whsec_... # absent -> webhooks are refused
scrapewright serve
| Endpoint | Purpose |
|---|---|
GET /v1/credits/packs | the price list — public, no key needed |
POST /v1/credits/checkout | start a purchase, returns a Stripe Checkout URL |
POST /v1/webhooks/stripe | payment notifications from Stripe |
The webhook endpoint takes no API key — Stripe is the caller, so the signature is the credential, and an unverified endpoint would be a free credit printer for anyone who guessed the URL. Three rules hold the integration up:
examples/stripe_smoke_test.py runs the whole path against Stripe's test mode
with the 4242 card. Any other provider plugs into the same two-method
BillingProvider protocol in scrapewright.service.billing; without one, the
service simply runs free, which is the right default for a demo or a self-hosted
instance.
Docker:
docker build -t scrapewright . # static paths
docker build -t scrapewright --build-arg WITH_JS=1 . # + headless Chromium
docker run -p 8000:8000 -v sw-data:/data scrapewright
Add --js (or Scrapewright(js=True)) and pages that render their catalog in the
browser become extractable:
scrapewright run https://spa-store.example.com/products/x --page --js
scrapewright crawl https://spa-store.example.com/shop --js -o products.xlsx
Rendering stays rare by construction: the static fetch runs first, and Chromium is
only started when the static HTML is an empty client-side shell or extraction on it
fails. A recipe learned from rendered HTML is tagged needs_js, so later runs on that
site skip the wasted static hop. The browser starts at most once per run and is reused
for every page.
Product shapeurl: str # canonical product URL
title: str
brand: str | None
price: Decimal | None # parsed from "1,250.00" / "1.250,00" / "€1290" alike
currency: str | None
available: bool | None
images: list[str] # absolute URLs
sizes: list[str]
description: str | None
sku: str | None
source_platform: str # shopify | woocommerce | json-ld | selector
A record is usable when it carries a title, a price, and a URL. The
validator (scrapewright.coverage) reports the usable ratio across a batch —
the number a recipe is trusted on before it's cached.
| Module | Role |
|---|---|
detect | Platform registry: free-catalog probes, then fingerprints for 12 platforms; returns the strategy to use |
extract/shopify, extract/woocommerce | Deterministic catalog extractors |
extract/jsonld | schema.org/Product from <script type="application/ld+json"> — free, ~common |
extract/llm | Synthesizes a SelectorRecipe from HTML — the one-time compile step |
extract/selectors | Replays a recipe with BeautifulSoup — the deterministic runtime |
schema | Schema/Field — declare what to extract; PRODUCT_SCHEMA is the built-in default |
service/ | FastAPI app: API keys (stored hashed), record-based quotas, cost metering, background crawl jobs, pluggable billing |
service/credits | Credit prices, packs, and the free allowance |
service/stripe_billing | Stripe Checkout + signature-verified webhook |
service/pricing | Measured unit costs and the margin each pack clears |
mcp_server | Five MCP tools so AI agents can call scrapewright directly |
fetch | StaticFetcher (plain HTTP) and BrowserFetcher (headless Chromium), plus the shell heuristic that decides when a render is worth paying for |
crawl | Frontier: turns one listing URL into product URLs (pattern match + card-template fallback + pagination) — deterministic, no LLM |
cache | Persists recipes keyed by domain, so the compile happens once |
validate | Field-coverage scoring |
export | Batch → .csv / .xlsx / .jsonl |
pipeline | Orchestrates detect → extract → validate → cache → heal |
max_synth_per_run (default 3) — a site that resists synthesis cannot burn
one model call per page. The bill is bounded no matter how large the crawl.model and works with any
injected client; the default targets Anthropic's Claude via the official SDK.The deterministic paths are fully covered by offline fixtures — no network, no model calls — so CI is green without an API key:
pip install "scrapewright[dev]"
pytest
v0.9 (alpha). Implemented and tested: an HTTP service with API keys, prepaid credits (priced off the one action that costs money, on an auditable ledger) and Stripe checkout with a signature-verified webhook, cost metering and background jobs; platform detection across 12 storefronts with a recommended strategy per site, catalog extraction (Shopify, WooCommerce), page extraction (JSON-LD, LLM-synthesized selectors), recipe caching, self-healing re-synthesis with a bounded per-run model budget, a crawl frontier (one listing URL → the whole site), JS rendering via an optional Playwright fetcher with automatic escalation, schema-agnostic extraction (bring your own fields), an MCP server for AI agents, coverage validation, and CSV / XLSX / JSONL export. 141 offline tests.
Known limit, stated plainly: it does not defeat anti-bot walls — deliberately out of scope. Sites behind Akamai/Fastly-style challenges return an honest miss.
Roadmap: pagination strategies for infinite-scroll listings, and a deployed instance of the service.
MIT — see LICENSE.
FAQs
Give it a store URL — it detects the platform and, for custom sites, synthesizes a reusable scraper once via an LLM, then replays it for free.
We found that scrapewright demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Security News
It has been one year since Shai-Hulud made its first appearance on npm.

Research
/Security News
Operators behind PolinRider used a compromised GitHub account to plant malware in four development versions of a Packagist package with 700,000+ downloads.

Security News
GitHub Actions now supports cache-mode, a least-privilege control on the Actions cache aimed at the cache poisoning technique behind recent compromises.