New:Microsoft Teams Notifications Are Now Available in Socket.Learn more →
Get Started

@three-ws/vision-mcp

Package Overview
Dependencies
Maintainers
1
Versions
3
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

@three-ws/vision-mcp

Let any AI agent see — analyze and describe images through the three.ws vision pipeline. Free NVIDIA NIM VLMs first, with an automatic server-side paid backstop. Pass an image URL or base64 and get a structured, natural-language read. Read-only, no key re

latest
Source
npmnpm
Version
0.1.2
Version published
Weekly downloads
58
-79.72%
Maintainers
1
Weekly downloads
 
Created
Source

three.ws

@three-ws/vision-mcp

Let any AI agent see — analyze and describe images through the three.ws vision pipeline, free-first and no key required.

npm license node MCP Registry three.ws

A Model Context Protocol server that gives any AI assistant image understanding over stdio. Pass an image as a public URL or base64 and the agent can read the text in a screenshot, identify what's in a scene, critique a 3D/avatar render, or write alt-text — all live, read-only, no key required.

Every answer comes from the three.ws vision pipeline: free NVIDIA NIM vision models lead every request, with an automatic server-side paid backstop behind them. The caller never pays and needs no provider key — point THREE_WS_BASE at a deployment and go.

Install

npm install @three-ws/vision-mcp

Or run with npx (no install):

npx @three-ws/vision-mcp

Quick start

Claude Code, one line:

claude mcp add vision -- npx -y @three-ws/vision-mcp

Claude Desktop / Cursor (claude_desktop_config.json or mcp.json):

{
	"mcpServers": {
		"vision": {
			"command": "npx",
			"args": ["-y", "@three-ws/vision-mcp"]
		}
	}
}

Inspect the surface with the MCP Inspector:

npx -y @modelcontextprotocol/inspector npx @three-ws/vision-mcp

Tools

ToolTypeWhat it does
analyze_imageread-onlyAnswer a prompt about an image — visual Q&A, OCR (read a screenshot), classification, or critiquing a render.
describe_imageread-onlyNatural-language description for alt text or a caption, steerable by detail level and focus.
get_vision_statusread-onlyReport whether a vision provider is live and which image formats are accepted. No image, no cost.

All three tools are read-only — they look at an image, they never store or mutate anything. A VLM is non-deterministic and the live provider chain moves, so none are idempotent.

Input parameters

analyze_image — prompt (required), and exactly one of imageUrl (public https) or image (base64 / data URI). Optional imageType (for raw base64), maxTokens (16–2048, default 512).

describe_image — exactly one of imageUrl or image. Optional imageType, detail (brief | standard | detailed, default standard), focus (an aspect to emphasize).

get_vision_status — no parameters.

Example

// analyze_image: read the text in a logo image
> { "imageUrl": "https://three.ws/logo.png", "prompt": "What does this logo say?" }
{
  "ok": true,
  "text": "The logo reads \"three.ws\".",
  "provider": "nvidia",
  "model": "nvidia/nemotron-nano-12b-v2-vl",
  "prompt": "What does this logo say?",
  "image_source": "url"
}
// describe_image — alt text from a base64 image
> { "image": "data:image/png;base64,iVBORw0KG...", "detail": "brief" }
{
  "ok": true,
  "description": "A seagull in profile against a blurred coastal background.",
  "detail": "brief",
  "provider": "nvidia",
  "model": "nvidia/nemotron-nano-12b-v2-vl",
  "image_source": "base64"
}
// get_vision_status — gate a "describe this" affordance
> {}
{
  "ok": true,
  "configured": true,
  "image_types": ["image/jpeg", "image/png", "image/webp", "image/gif"],
  "max_image_mb": 12
}

Examples

Runnable, key-free examples live in examples/:

node examples/list-tools.mjs     # every tool with its full input schema
node examples/read-an-image.mjs  # status + alt text + OCR on a live image

Both spawn this server over stdio and call the real vision pipeline. Every tool is read-only, so nothing is stored and nothing is paid. See examples/README.md for expected output.

Images may be JPEG, PNG, WebP, or GIF, up to 12 MB. Base64 inputs are size-checked locally before upload; imageUrl must be a public https URL (the vision server fetches it, so private/loopback hosts are rejected).

Pricing & auth

There is no charge to the caller and no key is needed. The platform runs every request against free NVIDIA NIM vision lanes first and falls back to a paid model only on its own infrastructure — that cost is the platform's, transparent to you.

The endpoint is rate-limited per IP. To raise your limit to the per-user tier, set the optional THREE_WS_API_KEY (a three.ws API key); it is sent as a bearer token. It is never required to use the free model.

Requirements

  • Node.js >= 20.
  • Network access to https://three.ws (or your own THREE_WS_BASE).

Environment variables

VariableRequiredDefaultPurpose
THREE_WS_BASEnohttps://three.wsWhich deployment to call.
THREE_WS_TIMEOUT_MSno30000Per-request timeout (a call may chain across free lanes).
THREE_WS_API_KEYno—Optional bearer token to raise the rate limit to per-user.

Part of the three.ws SDK suite — 3D AI agents, on-chain identity, and agent payments.
Website · Changelog · GitHub

Keywords

mcp

FAQs

Package last updated on 11 Sep 2026

Related posts