
Company News
Free Business Plan Upgrades for Open Source Maintainers
Open source maintainers are under more pressure than ever. We're raising our open source program from the Team plan to the Business plan, free.
@three-ws/vision-mcp
Advanced tools
Let any AI agent see — analyze and describe images through the three.ws vision pipeline. Free NVIDIA NIM VLMs first, with an automatic server-side paid backstop. Pass an image URL or base64 and get a structured, natural-language read. Read-only, no key re
Let any AI agent see — analyze and describe images through the three.ws vision pipeline, free-first and no key required.
A Model Context Protocol server that gives any AI assistant image understanding over stdio. Pass an image as a public URL or base64 and the agent can read the text in a screenshot, identify what's in a scene, critique a 3D/avatar render, or write alt-text — all live, read-only, no key required.
Every answer comes from the three.ws vision pipeline: free NVIDIA NIM vision models lead every request, with an automatic server-side paid backstop behind them. The caller never pays and needs no provider key — point THREE_WS_BASE at a deployment and go.
npm install @three-ws/vision-mcp
Or run with npx (no install):
npx @three-ws/vision-mcp
Claude Code, one line:
claude mcp add vision -- npx -y @three-ws/vision-mcp
Claude Desktop / Cursor (claude_desktop_config.json or mcp.json):
{
"mcpServers": {
"vision": {
"command": "npx",
"args": ["-y", "@three-ws/vision-mcp"]
}
}
}
Inspect the surface with the MCP Inspector:
npx -y @modelcontextprotocol/inspector npx @three-ws/vision-mcp
| Tool | Type | What it does |
|---|---|---|
analyze_image | read-only | Answer a prompt about an image — visual Q&A, OCR (read a screenshot), classification, or critiquing a render. |
describe_image | read-only | Natural-language description for alt text or a caption, steerable by detail level and focus. |
get_vision_status | read-only | Report whether a vision provider is live and which image formats are accepted. No image, no cost. |
All three tools are read-only — they look at an image, they never store or mutate anything. A VLM is non-deterministic and the live provider chain moves, so none are idempotent.
analyze_image — prompt (required), and exactly one of imageUrl (public https) or image (base64 / data URI). Optional imageType (for raw base64), maxTokens (16–2048, default 512).
describe_image — exactly one of imageUrl or image. Optional imageType, detail (brief | standard | detailed, default standard), focus (an aspect to emphasize).
get_vision_status — no parameters.
// analyze_image — read the text in a screenshot
> { "imageUrl": "https://three.ws/render.png", "prompt": "What does the banner text say?" }
{
"ok": true,
"text": "The banner reads \"three.ws — 3D AI agents\".",
"provider": "nvidia",
"model": "nvidia/nemotron-nano-12b-v2-vl",
"prompt": "What does the banner text say?",
"image_source": "url"
}
// describe_image — alt text from a base64 image
> { "image": "data:image/png;base64,iVBORw0KG...", "detail": "brief" }
{
"ok": true,
"description": "A seagull in profile against a blurred coastal background.",
"detail": "brief",
"provider": "nvidia",
"model": "nvidia/nemotron-nano-12b-v2-vl",
"image_source": "base64"
}
// get_vision_status — gate a "describe this" affordance
> {}
{
"ok": true,
"configured": true,
"image_types": ["image/jpeg", "image/png", "image/webp", "image/gif"],
"max_image_mb": 12
}
Images may be JPEG, PNG, WebP, or GIF, up to 12 MB. Base64 inputs are size-checked locally before upload; imageUrl must be a public https URL (the vision server fetches it, so private/loopback hosts are rejected).
There is no charge to the caller and no key is needed. The platform runs every request against free NVIDIA NIM vision lanes first and falls back to a paid model only on its own infrastructure — that cost is the platform's, transparent to you.
The endpoint is rate-limited per IP. To raise your limit to the per-user tier, set the optional THREE_WS_API_KEY (a three.ws API key); it is sent as a bearer token. It is never required to use the free model.
https://three.ws (or your own THREE_WS_BASE).| Variable | Required | Default | Purpose |
|---|---|---|---|
THREE_WS_BASE | no | https://three.ws | Which deployment to call. |
THREE_WS_TIMEOUT_MS | no | 30000 | Per-request timeout (a call may chain across free lanes). |
THREE_WS_API_KEY | no | — | Optional bearer token to raise the rate limit to per-user. |
Part of the three.ws SDK suite — 3D AI agents, on-chain identity, and agent payments.
Website · Changelog · GitHub
FAQs
Let any AI agent see — analyze and describe images through the three.ws vision pipeline. Free NVIDIA NIM VLMs first, with an automatic server-side paid backstop. Pass an image URL or base64 and get a structured, natural-language read. Read-only, no key re
We found that @three-ws/vision-mcp demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.
Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Company News
Open source maintainers are under more pressure than ever. We're raising our open source program from the Team plan to the Business plan, free.

Security News
The supply chain control that delays freshly published gems now covers lockfile generation and gem vendoring in Ruby projects.

Security News
During a UK cyber test, a Mythos 5 agent used sockpuppets, social engineering, and prompt injection to try to get a maintainer to merge malware.