New:Socket for Asana Is Now Available.Learn more
Get Started

dsh-tool-vision

Package Overview
Dependencies
Maintainers
1
Versions
18
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

dsh-tool-vision

DeepSeek Harness 外置视觉模型插件:inspect_image 把本地图片或 http(s) 图片 URL 发给任意 OpenAI 兼容端点,视觉模型看图的文字回答直接带回对话;附 Web UI 设置栏。

latest
Source
npmnpm
Version
0.6.4
Version published
Weekly downloads
480
-87.7%
Maintainers
1
Weekly downloads
 
Created
Source

dsh-tool-vision

中文文档

GitHub: Scorp1o117/dsh-tool-vision · npm: dsh-tool-vision

Enhancement Suite npm

Part of the DeepSeek Harness Enhancement Suite — Vision · Soul/Persona · Long-term Memory · Plugin Marketplace.

External vision model for DeepSeek Harness.

DSH 0.1.1 adds native image input for DeepSeek's vision catalog. This plugin remains useful when you want a separate OpenAI-compatible vision endpoint, pixel-level image tools, screenshots, or a text-model bridge. The harness derives every model request strictly from the session log (llm/stream requests must equal the durable derivation — the agent-loop invariant), so the bridge keeps its conversion inside that durable path:

  • inspect_image tool — sends an image (local file, or http(s) URL) to any OpenAI-compatible /chat/completions endpoint that supports image_url content parts, and returns the vision model's textual answer into the agent loop.
  • Image bridge (v0.2.1) — pasted images are turned into inspect_image hints before they enter the durable log, on the agent/pre-step waterfall (the one seam where the harness lets a plugin replace the messages of a proposed step). Images already logged by an older version are repaired lazily with a surface replace on the session's first pre-step. Only models listed in multimodalModels receive image blocks directly; a model's declared inputModalities are never consulted, because profiles routinely declare input: [text, image] on text-only models just to pass the harness's prompt-admission check.
  • Zero dependencies beyond the dsh SDK — works with any compatible endpoint: OpenAI GPT-4o, Qwen-VL (DashScope), GLM-4V (Zhipu), Moonshot, Gemini compatible endpoints, local Ollama, etc.
  • Registered on the global tools layer: every agent in the process can call inspect_image.
  • Web UI settings section (v0.3.0): Settings → 视觉模型 edits the tool-vision namespace (API endpoint, write-only key, model, bridge options) in settings.yaml; changes hot-apply without a restart. The API key lives in settings.yaml, not the profile patch. Mount by package name (name: 'dsh-tool-vision') so the web client bundle is discovered.

Install

Mount in a profile patch ($DSH_HOME/profiles/<name>/cordis.patch.yml):

- insert:
    - id: tool-vision
      name: 'dsh-tool-vision'     # after: pnpm add dsh-tool-vision in the profile
      config:
        baseURL: 'https://api.openai.com/v1'
        apiKeyEnv: 'VISION_API_KEY'
        model: 'gpt-4o-mini'

Or load it from a local path without npm:

    - id: tool-vision
      name: './plugins/dsh-tool-vision/index.js'

Config

FieldDefaultMeaning
baseURLhttps://api.openai.com/v1OpenAI-compatible API base URL.
apiKey''API key (takes precedence over env).
apiKeyEnvVISION_API_KEYEnv var holding the key.
modelgpt-4o-miniVision model id.
maxTokens1024Max output tokens.
timeoutMs60000Per-request timeout.
maxImageBytes10MBLargest accepted local image.
descriptiondefaultTool description shown to the model.
bridgeTextOnlytrueBridge pasted images to text hints on models that cannot see images.
bridgeExportDirtempExport dir for bridged images (os.tmpdir()/dsh-vision-bridge).
multimodalModels[]Model ids that receive image blocks directly (e.g. mimo-v2.5).
bridgePreviewtrueInline preview for bridged images: thumbnail above the hint text in the user bubble (click to zoom).
bridgePreviewScanIntervalMs2000Fallback scan interval for the preview scanner (ms); 0 disables the fallback.
bridgePreviewHideHinttrueHide the bridged hint text once the preview image has loaded (kept on failure — safe degradation).
bridgeAutoImagetrueWhile the bridge is on, report image input capability for every model to the host admission gate, so pasted images are accepted on text-only models without hand-editing provider configs.

Image bridge setup

  • (Optional, usually not needed) If bridgeAutoImage is disabled, declare image input on the models you paste images onto, so the harness admits image messages (pi-ai style):
    llm-pi-ai:
      providers:
        your-provider:
          models:
            - id: deepseek-v4-flash
              input: [text, image]
    
  • List genuinely multimodal models in the plugin config so they receive image blocks untouched:
    - id: tool-vision
      name: 'dsh-tool-vision'
      config:
        multimodalModels: ['mimo-v2.5', 'grok-4.5']
    

Then pasting an image while on a text-only model stores a hint like [User sent an image, exported to: <path>. Inspect it with the inspect_image tool...] in the transcript (the pasted image no longer renders as pixels in that message), and the agent inspects it through the configured vision endpoint.

Why not llm/stream? The harness freezes every request and the agent-loop invariant fails any request whose messages diverge from the session-log derivation (log-reconstruction desync), and this cordis waterfall's next() cannot replace request arguments. The agent/pre-step waterfall is the supported seam: its decision messages become the durable log, so the invariant stays satisfied.

Key resolution order: config.apiKeyprocess.env[apiKeyEnv]process.env.OPENAI_API_KEY.

Bridge image preview (v0.4.0)

On text-only models, pasted images become [User sent an image...] hint text in the transcript. With bridgePreview enabled (default), the browser half renders those hints as inline thumbnails in the display layer only:

  • Thumbnail + lightbox: click to zoom full-screen; click anywhere or press Esc to close;
  • Immediate + fallback: new messages are handled by a MutationObserver; history is back-filled by a periodic scan (interval via bridgePreviewScanIntervalMs);
  • Hide the hint (P2): with bridgePreviewHideHint on, the hint text is hidden once the image has loaded, leaving just the image; on load failure the text stays (safe degradation — never "no image AND no text");
  • Precise identification: bridged hints carry an invisible prefix marker (\u200b[bridge]), so ordinary user text that happens to contain "exported to:" is never misidentified;
  • Display-layer red line: persisted messages, the transcript, the model-facing text and the inspect_image chain are untouched.

Preview images are served by the same-origin loopback route /plugins/dsh-tool-vision/image: read-only access to the bridge export directory, localhost-only Host, image extensions only, ≤ 20MB per file, path-traversal protected.

Tool: inspect_image

ArgRequiredMeaning
pathImage path (absolute, or relative to the current workspace) or http(s) URL.
questionOptional specific question about the image.
detailauto / low / high resolution hint.

Example endpoints (baseURL):

  • OpenAI: https://api.openai.com/v1gpt-4o, gpt-4o-mini
  • Alibaba DashScope (Qwen-VL): https://dashscope.aliyuncs.com/compatible-mode/v1qwen-vl-plus, qwen-vl-max
  • Zhipu (GLM-4V): https://open.bigmodel.cn/api/paas/v4glm-4v-flash (free tier), glm-4v-plus
  • Moonshot (Kimi): https://api.moonshot.cn/v1moonshot-v1-8k-vision-preview
  • Ollama local: http://localhost:11434/v1llama3.2-vision (no key)

Note for users

  • This plugin is a standard profile bundle (dsh.bundle.patch): dsh plugin --profile web add dsh-tool-vision installs and mounts it in one step — no manual cordis.patch.yml edits needed.
  • Settings changes hot-apply (no restart needed).
  • Version 0.6.3 and newer require DSH 0.1.0-rc.7 or newer and are tested against 0.1.0-rc.7, 0.1.0-rc.8, and 0.1.1-rc.1.
  • DSH 0.1.0-rc.6 users must pin dsh-tool-vision@0.6.1, the last release carrying the legacy settings-allowlist compatibility patch.

Pixel-level vision tools (v0.6.0, ported from dsh-vision-router)

14 vision_* tools driven by the same configured endpoint as inspect_image (baseURL/apiKey/model) — no provider chain, no local models, no extra settings:

ToolPurpose
vision_describeImage Q&A / multi-image comparison (optional structured JSON)
vision_groundLocate a target and return its ORIGINAL-pixel bounding box
vision_detectEnumerate elements (buttons, inputs, icons…) with numbered boxes
vision_cropCrop a pixel region to a PNG artifact
vision_pixel_diffPer-pixel comparison: ratio, worst regions, heatmap, report
vision_colorsDominant-color quantization for palette matching
vision_ocrVerbatim text transcription (letters only — not scene analysis)
vision_long_screenshot_ocrChunked long-screenshot transcription into Markdown
vision_tracePotrace vectorization into colored SVG (worker-thread, safe)
vision_extract_foregroundSolid-background removal → transparent PNG
vision_html_screenshotHeadless render of a local .html (network blocked)
vision_screenshotDesktop capture (privacy-gated: enable desktopScreenshot in settings; Win: PowerShell / macOS: screencapture / Linux: import/scrot)
vision_presentPublish a generated image to the user via the host attachment store
vision_materializeCopy an attachment/local image into the workspace as a real path

Quality & safety details:

  • Content-hash cache keyed by endpoint+model+image+question (no stale answers across model switches, failures are never cached).
  • Uniform 4MP downscale before every model call; oversized inputs are rejected with a clear error (stat pre-check, 20MB cap on both file and attachment paths).
  • Rate-limit / 5xx auto-retry with Retry-After-aware backoff; endpoint content-safety rejections are surfaced as VISION_CONTENT_FILTERED instead of a generic backend error.
  • Long-OCR bounds: 120s total budget, 40-chunk cap, cancellation checks, stop-on-first-backend-failure.
  • Path containment for relative inputs; artifacts land in <workspace>/.dsh-tool-vision/.

Requires sharp / potrace / puppeteer-core (declared as optional dependencies: a failed platform install never blocks the plugin; missing ones degrade lazily with an install hint and never break other tools).

vision_screenshot is privacy-sensitive and therefore not registered by default — set desktopScreenshot: true in the tool-vision settings to enable desktop capture.

Limitations

  • A bridged image enters the conversation as a text hint (a transcript, not pixels) — pixel-precise in-context reasoning is not available to text-only models; the vision model's description comes back through inspect_image.
  • Images are base64-transferred; mind privacy and size limits.
  • Independent of the dsh-llm routing/retry system; failures return clear errors to the agent.

License

MIT — bridge preview & integration: xing666173. Pixel vision tools ported from dsh-vision-router (© ysr666, MIT) with gratitude.

Keywords

dsh

FAQs

Package last updated on 24 Aug 2026

Related posts