New:Microsoft Teams Notifications Are Now Available in Socket.Learn more
Get Started

pyrecrawl

Package Overview
Dependencies
Maintainers
1
Versions
11
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

pyrecrawl

Web browsing superpowers for AI agents: one MCP tool to scrape, extract, crawl, map, and search any site — self-hosted, no API keys, with a smart auto-fallback ladder.

pipPyPI
Version
0.8.0
Weekly downloads
803
190.94%
Maintainers
1
Weekly downloads
 
Created

🔥 PyreCrawl — Web Browsing Superpowers for Your AI Agent

License: MIT MCP Python 3.10+ PyPI GitHub stars Downloads / 30d Active users / 30d

One command gives any AI agent the whole web. Scrape, extract, crawl, map, and search — self-hosted, no API keys, no rate limits, no subscription.

PyreCrawl speaks MCP (Model Context Protocol), the standard tool interface for Claude, Cursor, VS Code, Codex, OpenCode, Hermes, and any MCP-compatible agent.

A smart auto-fallback ladder always picks the cheapest method that succeeds:

fast HTTP
    │  (403/503/Cloudflare challenge or empty body)
    ▼
stealth browser (real Chromium + Cloudflare solver)
    │  (still blocked, or the page needs full JS rendering)
    ▼
deep processing (LLM-ready markdown, citations, structured extraction)

⚡ Tools exposed

ToolWhat it does
scrape(url, prefer="auto")Single URL → LLM-ready markdown
extract(url, schema)Scrape + structured extraction (JsonCss schema)
map_site(root, include_pattern=None, limit=200)Enumerate all internal URLs
crawl(root, max_pages=5, prefer="auto", include_paths=None, exclude_paths=None, max_depth=0)Multi-page crawl with path filters + true BFS depth
document(url)PDF/DOCX/PPTX → markdown (no browser, optional [docs] extras)
search(query, limit=10)Web search via DuckDuckGo HTML (no API key)
search_papers(query, limit=8, source="arxiv", category=None)Academic search via arXiv + Crossref (no API key) — feed pdf_url into document
batch_scrape(urls[], ...)Many URLs in ONE call — parallel, deduped, cache-aware
deep_research(query, limit=5, scrape_top=3)Search → evidence pack with [n] citations (no LLM synthesis — your agent does that)
monitor(url, action, css_selector=None)Change detection with persisted snapshots + unified diff
session(session, action, ...)Persistent browser session (cookies kept) — login walls, multi-step flows, screenshots
cache(action)Inspect/clear/enable/disable the HTTP response cache
health()Versions + import sanity check

MCP Resources (read-only state without a tool call): pyrecrawl://cache/stats · pyrecrawl://sessions · pyrecrawl://monitors

MCP Prompts (ready-made playbooks): research(topic) · rag_ingest(site) · watch_page(url)

Env flags

VariableDefaultEffect
PYRECRAWL_CACHEoff1 = in-memory LRU (128 pages), or a directory path (reserved for disk mode)
PYRECRAWL_CACHE_TTL900Cache entry lifetime in seconds
PYRECRAWL_MONITOR_DIR~/.pyrecrawl/monitorsWhere monitor snapshots persist
PYRECRAWL_NO_TELEMETRYoff1 = disable the anonymous startup ping (also honors DO_NOT_TRACK=1)

prefer options: "auto" (default ladder) · "fast" (HTTP only) · "stealth" (CF bypass) · "llm" (deep processing).

🚀 Install & Use (one-liner)

1. Install

UV is a fast Python package manager that handles Python itself — no need to install Python separately. Get it once:

# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh

# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

Learn more about UV →

Then run PyreCrawl directly — no venv, no pip install, no Python download:

uvx pyrecrawl@latest
uv tool install pyrecrawl

Or via pipx (alternative)

pipx install pyrecrawl

Or via pip into a venv

pip install pyrecrawl

2. One-time browser engines

pyrecrawl setup

This installs Chromium + stealth browser engines (~2 min, one-time).

3. Register with your AI agent

# Auto-detect installed agents and write their MCP configs
pyrecrawl install

# Or target specific agents
pyrecrawl install claude-desktop cursor

# Dry-run to preview what would change
pyrecrawl install --dry-run

Supported agents: claude-desktop, claude-code, cursor, vscode, codex, opencode, hermes.

4. Start chatting

After installing + registering, restart your agent (or start a new session). Then ask:

"Scrape https://example.com and summarize it."

Available tools:

ToolWhat it does
scrapeFetch a single URL → markdown (auto-escalates past Cloudflare)
extractScrape + structured extraction via CSS schema → JSON
map_siteEnumerate all internal URLs from a root
crawlMulti-page crawl: discover + scrape in bulk
batch_scrapeFetch many URLs in one parallel call
searchWeb search via DuckDuckGo with anti-bot bypass
search_papersAcademic paper search (arXiv / Crossref)
deep_researchSearch + scrape + citations in one call — primary research tool
documentExtract text from PDF/DOCX/PPTX URLs
monitorTrack a URL for content changes over time
sessionPersistent browser session for login walls
cacheInspect or clear the response cache
healthVerify engine availability + version

Plus 3 guided prompts: research, rag_ingest, watch_page.

Quick examples

Ask your agent naturally — no special syntax needed:

You sayAgent uses
"Scrape https://example.com and summarize it"scrape → returns markdown → agent summarizes
"Research Rust memory safety vulnerabilities"deep_research → search + scrape + citations
"Deep research on AI regulation worldwide"deep_research(iterations=3) → multi-pass with refined queries
"Extract all product names and prices from this page"extract → CSS schema → structured JSON
"Crawl https://docs.example.com and give me an overview"crawl → multi-page → summary
"Monitor this page for price changes"monitor → baseline snapshot → periodic diff
"Find papers about transformer attention"search_papers → arXiv results
"What's the current cache hit rate?"cache → stats

🧠 Skills — Maximize Your Agent's Research Quality

PyreCrawl tools give your agent hands (scrape, crawl, search). But the agent still needs a brain — instructions on when to use which tool, how to chain research passes, and what anti-hallucination rules to follow.

That's what PyreCrawl Skills provides.

MCP Tools (this repo)Skills (pyrecrawl-skills)
RoleExecute web operationsTell the agent how to use them
AnalogyHandsBrain
Exampledeep_research(query, iterations=3)"Run 3 passes, check gaps after each, cite everything"
Required?Yes (the engine)Optional (but recommended for research quality)

Quick setup:

# 1. Install the tools (you already have this)
uvx pyrecrawl@latest

# 2. Add the research skill to your project
git clone https://github.com/SanggonBoy/pyrecrawl-skills.git /tmp/pyrecrawl-skills
cp /tmp/pyrecrawl-skills/pyrecrawl-research/SKILL.md ./CLAUDE.md  # or .cursorrules / AGENTS.md

Without skills: Your agent has powerful tools but improvises usage. With skills: Your agent follows a proven research protocol with anti-hallucination guardrails.

📚 Manual config (if pyrecrawl install doesn't match your setup)

Claude Desktop

Config file

  • Linux: ~/.config/Claude/claude_desktop_config.json
  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %AppData%\Claude\claude_desktop_config.json
{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Claude Code

Config file: project-scoped .mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Cursor

Config file: ~/.cursor/mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

VS Code / Copilot

Config file: .vscode/mcp.json (project-scoped)

{
  "servers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"],
      "type": "stdio"
    }
  }
}

Codex CLI

Config file: ~/.codex/config.toml

[mcp_servers.pyrecrawl]
command = "uvx"
args = ["--from", "pyrecrawl", "pyrecrawl", "serve"]

OpenCode

Config file: ~/.config/opencode/opencode.json

{
  "mcp": {
    "pyrecrawl": {
      "type": "local",
      "command": ["uvx", "--from", "pyrecrawl", "pyrecrawl", "serve"],
      "enabled": true
    }
  }
}

Hermes

Config file

  • Linux/macOS: ~/.hermes/config.yaml
  • Windows: %LocalAppData%\hermes\config.yaml
mcp_servers:
  pyrecrawl:
    command: uvx
    args:
      - --from
      - pyrecrawl
      - pyrecrawl
      - serve
    enabled: true

Windows note: uvx must be on PATH. If not, use the full path to uvx.exe (e.g. C:\Users\<you>\AppData\Local\hermes\bin\uvx.exe).

🧠 How the ladder chooses

PyreCrawl runs each request through three tiers, stopping at the first one that returns a complete, LLM-ready result:

ConcernFast tierStealth tierDeep tier
Static HTML page✅ ~200ms
Cloudflare-protected✅ Turnstile solver
JS-heavy SPA✅ real Chromium
Live DOM data (input .value, JS state)js param
LLM-ready markdown + citations✅ BM25, fit-markdown
Structured extraction (CSS schema)
Deep crawl (BFS/DFS/BestFirst)✅ adaptive

The agent never has to pick. prefer="auto" does it every call.

Live DOM data with js and wait_for

Some sites keep the data you want in a DOM property (e.g. an <input>'s .value) that JS writes after an XHR — it never appears in the serialized HTML. The scrape tool accepts two stealth-tier params for exactly this:

{
  "url": "https://temp-mail.org/id",
  "prefer": "stealth",
  "wait_for": "document.getElementById('mail').value.includes('@')",
  "js": "document.getElementById('mail').value"
}
  • wait_for — a JS predicate expression polled until truthy (bounded by timeout). Use it instead of guessing a sleep for anything that arrives asynchronously.
  • js — a JS expression evaluated once the page settles; the value comes back in meta.js_result. Errors are captured in meta.js_error (the page result is still returned, never a crash).

📊 Compared to Firecrawl (hosted)

FirecrawlPyreCrawl
CostFree 1k/mo, then $16–333/moFree, self-hosted
Local LLM support✅ Ollama / any LLM
Cloudflare bypass✅ (Fire-Engine, paid)✅ (free, built-in)
Markdown + BM25
Self-host
Academic paper search✅ arXiv + Crossref (search_papers)
Hosted search API✅ /search⚠️ DuckDuckGo HTML + arXiv/Crossref (no key)

🔧 Development

git clone https://github.com/SanggonBoy/PyreCrawl.git
cd PyreCrawl
uv venv --python 3.12 .venv
source .venv/Scripts/activate  # Windows; or .venv/bin/activate on macOS/Linux
uv pip install -e ".[dev]"
python -m playwright install chromium
scrapling install

Run tests

python scripts/selfcheck.py        # real-network smoke test (13 tools + engines)
python scripts/probe_stdio.py      # stdio JSON-RPC probe
python scripts/test_ladder_bug.py  # SPA-shell ladder escalation regression
python scripts/test_js_eval.py     # stealth js/wait_for params regression
python scripts/test_scope_selector.py  # crawl css_selector/max_depth wiring
python scripts/test_link_harvest.py    # map/BFS link purity regression

📦 Publish

Maintainers only:

git tag vX.Y.Z
git push origin vX.Y.Z

GitHub Actions builds + uploads to PyPI via trusted publishing.

🔔 Stay up to date

PyreCrawl checks PyPI on every startup and reports the latest version — your MCP agent sees this automatically via the health() tool response and can notify you inline.

To check manually:

pyrecrawl version

To upgrade:

pyrecrawl update   # runs: uv tool upgrade pyrecrawl

Get notified of new releases: click WatchReleases only at the GitHub repo to receive email notifications when a new version is published.

[!NOTE] PyreCrawl sends one anonymous usage ping per 24 h at server startup — see Privacy for exactly what's sent and how to opt out.

🔒 Privacy — anonymous usage ping

PyreCrawl phones home once per 24 h with a tiny anonymous ping when the MCP server starts, so we can count real users (DAU/MAU) instead of raw downloads.

Sent (4 fields, ~100 bytes)Never sent
Hashed machine id (SHA-256 of hostname+MAC — not reversible)Your IP (not stored)
PyreCrawl versionAny URL you scrape
Python versionAny page content or search queries
OS family (windows / linux / darwin)Anything else

Client code: src/pyrecrawl/telemetry.py (~90 lines, stdlib only) · Collector: workers/telemetry/ — a self-hostable Cloudflare Worker + D1, no third-party analytics service.

Opt out any time:

export PYRECRAWL_NO_TELEMETRY=1   # or the industry-standard DO_NOT_TRACK=1

📜 Uninstall

# Remove from all agent configs
pyrecrawl uninstall

# Remove the package
uv tool uninstall pyrecrawl

🛡️ License

MIT — see LICENSE.

Keywords

ai-agent

FAQs

Related posts