🎩 You're Invited:Meet the Socket team at Black Hat in Las Vegas, August 3-6.RSVP
Sign In

@sylphx/video-reader-mcp

Package Overview
Dependencies
Maintainers
2
Versions
6
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

@sylphx/video-reader-mcp

Cue — Evidence-first video reading for AI agents — ffprobe, subtitles, scenes, transcripts, and timelines without frame-by-frame LLM vision.

latest
Source
npmnpm
Version
0.1.7
Version published
Maintainers
2
Created
Source

Cue

Timeline proof for agents

Cue (transitional package @sylphx/video-reader-mcp) — local video timeline evidence, not frame-by-frame VLM.

video-reader-mcp — Sylphx Mark banner

Your agent watched the video. Did it read the timeline?

Evidence-first video reading for AI agents. One call turns any local video into a timeline document — ffprobe streams, chapters, embedded subtitles, scene boundaries, and warnings you can cite without frame-by-frame vision LLM calls.

npm version License CI/CD TypeScript

Local-first · One smart read_video call · Timeline evidence + provenance · 20 tests

SOTA family roadmap: docs/roadmap/sota-family-roadmap.md.

⭐ Star this repo if agents should read video timelines with proof, not sampled frame captions. · Quick start · See it work · Why not frame-by-frame vision?

This repository is product SSOT. Sibling agent tools live in separate repos. Each instrument is an independent repository (marketplace + stars).

Read video structure (not N-second frame spam)

Cue is local-first timeline architecture: streams, dialogue, scene cuts, structural keyframes, and agent_index for text-only agents.

Spec: docs/specs/agent-video-read-contract.md

Local-first frontier: ffmpeg/ffprobe + structural keyframes, optional local whisper ASR. No cloud required.

Product docs

DocPurpose
docs/POSITIONING.mdStrategic positioning
docs/COMPETITIVE.mdPeer anchors and wedge
docs/EVIDENCE_CONTRACT.mdEvidence = result contract
docs/TOOL_SURFACE.mdFew clear tools policy
docs/PRODUCT_INDEPENDENCE.mdThis repo is SSOT
docs/IPPB.mdIndependent public product bar
docs/PUBLISH.mdnpm/git publish status

The problem

Videos are not a pile of frames. They are codecs, chapters, embedded subtitles, scene cuts, variable frame rates, and timelines measured in milliseconds.

Most agent stacks sample frames and ask a vision LLM what it sees. Subtitles get skipped. Scene boundaries vanish. Duration and stream metadata never reach context. Citations become "around minute two, I think." Then the agent hallucinates — confidently.

Video Reader MCP is built for the moment your agent needs a citeable timeline, not a slideshow summary.

Why not frame-by-frame vision?

Typical vision pathVideo Reader MCP
Sample N frames into a vision modelffprobe format + stream metadata in one call
Paraphrased "what happens"Embedded subtitle cues with start_ms, end_ms, and provenance
Scene changes guessed from captionsOptional ffmpeg scene filter with timestamp evidence
Missing audio / VFR silently ignoredWarnings for missing ffmpeg/ffprobe, VFR, missing audio, skipped ASR
Cloud API by defaultLocal-first — ffprobe + ffmpeg on your machine
Ship and pray20 tests on parsers, fixture corpus, doctor, release gate, and integration

See it work

Install (30 seconds)

npm install -g @sylphx/video-reader-mcp
cue doctor
claude mcp add cue -- npx @sylphx/video-reader-mcp

Install once. Call once.

claude mcp add video-reader -- npx @sylphx/video-reader-mcp
{
  "sources": [{ "path": "/absolute/path/to/demo.mp4" }],
  "include_subtitles": true,
  "include_scenes": true
}

read_video builds a timeline document per source — no per-frame vision LLM calls:

{
  "source": "/absolute/path/to/demo.mp4",
  "success": true,
  "data": {
    "provenance": {
      "source": "/absolute/path/to/demo.mp4",
      "tool": "read_video",
      "version": "0.1.0",
      "extracted_at": "2026-07-09T12:00:00.000Z"
    },
    "format": {
      "format_name": "mov,mp4,m4a,3gp,3g2,mj2",
      "duration_ms": 125500
    },
    "streams": [
      { "index": 0, "codec_type": "video", "width": 1920, "height": 1080 },
      { "index": 1, "codec_type": "audio", "channels": 2, "sample_rate": 48000 }
    ],
    "chapters": [
      { "id": 0, "start_ms": 0, "end_ms": 60250, "title": "Intro" }
    ],
    "subtitles": [
      {
        "index": 0,
        "start_ms": 1200,
        "end_ms": 3400,
        "text": "Welcome to the demo.",
        "provenance": { "method": "ffmpeg_extract", "format": "srt" }
      }
    ],
    "scenes": [
      {
        "index": 0,
        "time_ms": 45200,
        "provenance": { "method": "ffmpeg_scene_filter", "threshold": 0.4 }
      }
    ],
    "warnings": []
  }
}

Abbreviated shape — optional local ASR transcript hooks skip gracefully when no adapter is wired.

Prerequisites

  • Node.js >=22.13
  • ffprobe (required) and ffmpeg (recommended for subtitles + scenes) on PATH

MCP Tool Surface

ToolUse it when the agent needs to...
read_videoRead one or more local videos and return ffprobe metadata, chapters, subtitles, scenes, and timeline warnings.

Supported formats: MP4, M4V, MKV, MOV, WebM, and other formats ffprobe can inspect.

Quick Start

Claude Code

claude mcp add video-reader -- npx @sylphx/video-reader-mcp

Claude Desktop

Add this to claude_desktop_config.json:

{
  "mcpServers": {
    "video-reader": {
      "command": "npx",
      "args": ["@sylphx/video-reader-mcp"]
    }
  }
}

Any MCP Client

npx @sylphx/video-reader-mcp

HTTP transport (optional)

MCP_TRANSPORT=http MCP_HTTP_PORT=8080 npx @sylphx/video-reader-mcp

Security model

  • Local-firstread_video inspects local files; remote URLs are not fetched by default.
  • ffprobe/ffmpeg boundary — probe and frame tools shell out to configured binaries on PATH; missing tools return explicit errors.
  • Fixture corpus — CI validates parser and safety fixtures; corrupted inputs fail closed with structured diagnostics.
  • Evidence envelope — timestamps, frame indices, and extraction routes are preserved so agents can verify claims.

Release proof

Claims are backed by CI benchmark:release-gate, fixture corpus checks, and the shipped-path matrix (Rust-default primary tools).

bun run benchmark:release-gate

Artifact: benchmark-artifacts/video_reader_release_gate.json — must report status: passed before release.

Development

git clone https://github.com/SylphxAI/video-reader-mcp.git
cd video-reader-mcp
bun install
bun run build
bun test
bun run doctor
bun run benchmark:release-gate

Useful checks:

bun run check
bun run typecheck
bun run benchmark:release-gate

Example read_video requests live in examples/. CI runs parser, fixture corpus, doctor, and release-gate checks; integration tests exercise ffmpeg when available on the runner.

Support

Help this reach more builders

If frame-by-frame vision guesses have wasted your context, your citations, or your trust in agent output, you are exactly who this project is for.

⭐ Star the repo — it is the fastest way to help more agent builders find evidence-first video reading. Share it in your MCP client setup, team wiki, or agent stack README.

Discovery (in progress)

ChannelStatus
Glama MCP directoryListed — claim server for full discoverability
Official MCP RegistryListed — io.github.SylphxAI/video-reader-mcp @ v0.1.0
TensorBlock MCP Index PR #1113Open — multimedia/document processing listing
MCP servers community issue #4500Open — community server highlight
mcp.so listing issue #3068Open — directory submission request
mcpservers.org submitNot listed yet — free web-form submission

Know another MCP directory? Open an issue with the link.

License

MIT © SylphxAI

Keywords

mcp

FAQs

Package last updated on 01 Aug 2026

Did you know?

Socket

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Install

Related posts