🎩 You're Invited:Meet the Socket team at Black Hat in Las Vegas, August 3-6.RSVP
Sign In

pdf-insight-mcp

Package Overview
Dependencies
Maintainers
1
Versions
2
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

pdf-insight-mcp

MCP server for extracting text, images, tables, links, annotations, and metadata from PDF files

pipPyPI
Version
0.2.1
Weekly downloads
76
Maintainers
1

pdf-reader-mcp

一个功能丰富的 PDF 阅读 MCP 服务器,让 LLM(大语言模型)客户端能够读取和分析 PDF 文件。
A feature-rich MCP server for reading and analyzing PDF files with LLM clients.

功能特性 / Features

工具 / Tool中文说明English
get_pdf_info读取文档元数据、页数、大小和加密状态Read document metadata, page count, size, and encryption status
read_pdf_as_text提取指定页面文本内容Extract text content from selected pages
read_pdf_as_images将指定页面渲染为 base64 图片Render selected pages as base64-encoded images
get_pdf_outline读取书签与目录结构Read bookmarks and outline structure
search_pdf_text按页返回搜索结果和上下文Search text with per-page context
extract_pdf_tables提取可识别的表格结构Extract structured tables when detectable
extract_pdf_images提取 PDF 内嵌图片Extract embedded images from the PDF
get_pdf_page_info查看单页尺寸、文本、图片和链接信息Inspect a page's dimensions, text, images, and links
extract_pdf_links提取外部链接和内部跳转Extract external URLs and internal page jumps
get_pdf_annotations读取批注、高亮与注释信息Read comments, highlights, and annotation data
get_pdf_text_stats统计文本、行数、段落数和扫描版概率Compute text, line, paragraph, and scan-likelihood stats
compare_pdf_pages比较两个页面的文本相似度Compare text similarity between two pages

为什么做这个项目 / Why this project

很多 LLM 工作流不仅需要纯文本提取,还需要目录、表格、图片、注释、链接等结构化信息。
Many LLM workflows need more than raw text extraction. They also need structure, tables, images, annotations, and links.

这个服务提供统一的 MCP 接口,用于: This server provides a unified MCP interface for:

  • 文本型 PDF / text-heavy PDFs
  • 扫描版或版式敏感 PDF / scanned or layout-sensitive PDFs
  • 表格与图片提取 / table and image extraction
  • 元数据与结构分析 / metadata and structure inspection
  • 批注与链接分析 / annotation and link analysis

安装 / Installation

前置要求 / Prerequisites

  • Python 3.10+
  • uv 或其他 Python 环境管理工具 / uv or another Python environment manager

安装 uv / Install uv:

curl -LsSf https://astral.sh/uv/install.sh | sh

Windows PowerShell:

irm https://astral.sh/uv/install.ps1 | iex

从 PyPI 安装 / Install from PyPI

发布后可直接通过 uvx 运行: After the package is published, you can run it directly with uvx:

uvx pdf-insight-mcp

也可以先安装再运行: You can also install first, then run:

python -m pip install pdf-insight-mcp
pdf-reader-mcp

本地开发安装 / Local development setup

uv sync

运行服务 / Run the server

uv run pdf-reader-mcp

在 MCP 客户端中配置 / Configure in an MCP client

PyPI 安装方式示例 / Example config using the published PyPI package:

{
  "mcpServers": {
    "pdf-reader": {
      "command": "uvx",
      "args": ["pdf-insight-mcp"]
    }
  }
}

本地仓库开发配置示例 / Example configuration for a local checkout:

{
  "mcpServers": {
    "pdf-reader": {
      "command": "uv",
      "args": [
        "--directory",
        "/absolute/path/to/pdf-reader-mcp",
        "run",
        "pdf-reader-mcp"
      ]
    }
  }
}

/absolute/path/to/pdf-reader-mcp 替换为你的本地仓库路径。
Replace /absolute/path/to/pdf-reader-mcp with your local repository path.

发布 / Release

推荐发布路径: Recommended release path:

  • 发布 Python 包到 PyPI / Publish the Python package to PyPI
  • 发布 server.json 到官方 MCP Registry / Publish server.json to the official MCP Registry

建议使用 GitHub Actions + PyPI Trusted Publishing(OIDC)+ MCP Registry GitHub OIDC。 The recommended automation is GitHub Actions + PyPI Trusted Publishing (OIDC) + MCP Registry GitHub OIDC.

典型发布流程: Typical release flow:

# 1. 修改版本号(pyproject.toml 和 server.json)
# 2. 提交改动
git commit -am "Release v0.2.0"

# 3. 打 tag
git tag v0.2.0

# 4. 推送分支和 tag
git push origin main --tags

工作流会在 v* tag 上: The release workflow will, on v* tags:

  • 运行测试 / run tests
  • 构建 sdist 和 wheel / build sdist and wheel
  • twine check / run twine check
  • 发布到 PyPI / publish to PyPI
  • 发布到 MCP Registry / publish to the MCP Registry

响应大小与大 PDF 注意事项 / Response size and large-PDF notes

  • read_pdf_as_images 返回的是 base64 图片,响应体积会迅速变大。
    read_pdf_as_images returns base64 image payloads, which can grow very quickly.
  • 图片渲染仍然限制为最多 20 页。
    Image rendering is still limited to 20 pages per call.
  • read_pdf_as_text 现在默认限制为最多 50 页、最多 200000 字符,超限会截断并附带 warning。
    read_pdf_as_text now defaults to at most 50 pages and 200000 characters, and truncates with a warning when needed.
  • read_pdf_as_images 现在默认限制总返回负载约 20MB,超限会提前停止并附带 warning。
    read_pdf_as_images now defaults to an overall payload cap of about 20MB and stops early with a warning.
  • 对扫描版 PDF,建议优先按小页范围调用,并降低 dpi、使用 jpeg、降低 quality
    For scanned PDFs, prefer smaller page ranges, lower dpi, jpeg, and lower quality.

开发 / Development

安装开发依赖 / Install dev dependencies:

uv sync --extra dev

运行测试 / Run tests:

uv run pytest

技术栈 / Tech stack

  • Python 3.10+
  • MCP Python SDK
  • PyMuPDF

License

MIT

Keywords

llm

FAQs

Did you know?

Socket

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Install

Related posts