
Product
Socket Now Protects the Microsoft Edge Extension Ecosystem
Enterprise security teams can now detect malware, credential theft, suspicious network activity, and risky updates across Microsoft Edge extensions.
dsh-vision-recognizer
Advanced tools
Adaptive image routing for DeepSeek Harness: pass images directly to native multimodal models and transcribe them only for text-only models.
English | 简体中文
Keep DeepSeek as the conversation brain, attach images anyway, and switch the image-recognition provider any time from Settings → Plugins. A vision plugin for DeepSeek Harness.
It registers an adaptive provider route (default vision-recognizer, shown as DeepSeek + 智能识图 in the model picker) that wraps the configured conversation provider. The wrapper always admits image attachments, then resolves the exact selected model: models declaring native image input receive the original image blocks directly; text-only or unknown-capability models receive text transcribed by the vision model you configure. DeepSeek remains the default wrapped conversation provider.
attached image ──▶ vision-recognizer route ──▶ selected model supports image? ── yes ─▶ native image request
│
no
▼
configured vision transcription ──▶ text-only selected model
dsh plugin --profile web add dsh-vision-recognizer — no build scripts, no sharp approval (no native dependencies at all)./chat/completions) and native Anthropic Messages — Claude works out of the box.fallbackModels entry is tried in order (each may target a different vendor); only after all fail does the request fail, listing every attempt.autoLocalOllama (default on) probes http://localhost:11434 and prepends a running Ollama to the chain — images never leave your machine.dsh plugin --profile web add dsh-vision-recognizer
Slow npm registry?
dsh plugin --profile web add dsh-vision-recognizer --registry=https://registry.npmmirror.com
Install from a local checkout (development):
dsh plugin --profile web add file:/path/to/dsh-vision-recognizer
Use the
file:prefix (copies the package intonode_modules). A bareadd .oradd link:…makes pnpm symlink the package, in which case the plugin'sschemasterydependency resolves from the source checkout and is not found — a general pnpm symlink-install gotcha, not a bug in the plugin.
Restart dsh web, then:
[图片转译] result.With a native multimodal selected model, no fallback key is required. With a text-only model and no key or local Ollama, the turn fails fast with guidance instead of hanging.
Scope: adaptive fallback applies while the DeepSeek + 智能识图 wrapper route is selected. Selecting another provider route calls that route directly. rc8 does not expose a public decorator hook that can add fallback behavior to every existing provider route.
| Provider | baseURL | Default model | Key env var | Protocol |
|---|---|---|---|---|
| OpenAI | https://api.openai.com/v1 | gpt-4o-mini | OPENAI_API_KEY | OpenAI |
| Anthropic Claude | https://api.anthropic.com/v1 | claude-3-5-sonnet-latest | ANTHROPIC_API_KEY | Anthropic |
| Google Gemini | https://generativelanguage.googleapis.com/v1beta/openai | gemini-2.0-flash | GEMINI_API_KEY | OpenAI |
| OpenRouter | https://openrouter.ai/api/v1 | qwen/qwen-2.5-vl-72b-instruct | OPENROUTER_API_KEY | OpenAI |
| Azure OpenAI | user-supplied (…/openai/deployments/<deployment>) | gpt-4o-mini | AZURE_OPENAI_API_KEY | OpenAI |
| Ollama (local) | http://localhost:11434/v1 | auto-detected | none | OpenAI |
| Alibaba DashScope | https://dashscope.aliyuncs.com/compatible-mode/v1 | qwen-vl-max | DASHSCOPE_API_KEY | OpenAI |
| QwenCloud (Intl) | https://dashscope-intl.aliyuncs.com/compatible-mode/v1 | qwen-vl-plus | DASHSCOPE_API_KEY | OpenAI |
| Zhipu GLM | https://open.bigmodel.cn/api/paas/v4 | glm-4v-flash | ZHIPU_API_KEY | OpenAI |
| Baidu Qianfan | https://qianfan.baidubce.com/v2 | ernie-4.5-vl-8k | QIANFAN_API_KEY | OpenAI |
| iFlytek Spark | https://spark-api-open.xf-yun.com/v1 | generalv3.5 | SPARK_API_KEY | OpenAI |
| Moonshot Kimi | https://api.moonshot.cn/v1 | moonshot-v1-8k-vision-preview | MOONSHOT_API_KEY | OpenAI |
| Tencent Hunyuan | https://api.hunyuan.cloud.tencent.com/v1 | hunyuan-vision | HUNYUAN_API_KEY | OpenAI |
| Volcengine Doubao | https://ark.cn-beijing.volces.com/api/v3 | doubao-1.5-vision-pro-32k-250115 | ARK_API_KEY | OpenAI |
| SiliconFlow | https://api.siliconflow.cn/v1 | Qwen/Qwen2.5-VL-72B-Instruct | SILICONFLOW_API_KEY | OpenAI |
Model ids drift over time; the defaults are starting points — override
Modelin the settings UI. Key resolution order: key entered in the UI → the provider env var →$VISION_API_KEY/$DASHSCOPE_API_KEY.
Config saved from the UI is written to $DSH_HOME/vision-recognizer.json and merged over the bundle defaults at startup. cordis.patch.yml only carries factory defaults; a user cordis.patch.yml override still works as the composition-time fallback.
⚠️ patch semantics: the bundle's
- insert:appends this row to the entry list. Writing a second- insert:with the same id in your owncordis.patch.ymlwould register the adapter twice (undefined behavior). To override individual keys, write a single top-level- id: dsh-vision-recognizerentry; better yet, use the Settings UI.
The adaptive wrapper uses rc8 public interfaces only:
ctx.llm.listModels(innerProvider) and resolveModelInfo(innerProvider, model) inspect the exact target model and rebind its metadata to the wrapper route;ctx.llm.prepareCall(...) delegates to the configured target provider without depending on private adapter registrations;resolveModel advertises ['text', 'image'] so the wrapper admits images, while proxy stream uses the target's original modality declaration to choose native pass-through or transcription;settings.plugins.tab slot plus custom webServer routes, persisting config to its own JSON file.Capability lookup and prepared target dispatch are separate public operations in rc8. A target adapter replaced by HMR in that tiny interval can race the routing decision. Nested target delegation also enters the llm/stream waterfall a second time, and DSH may strip provider-private replay metadata when wrapper and target adapters differ. Ordinary text/image history is preserved; provider-specific replay signatures may lose their optimization or fidelity until DSH exposes an atomic delegation handle.
Native multimodal routing sends image bytes to the selected conversation provider. The text-only fallback instead sends them (base64, normally HTTPS) to the vision endpoint you configure. In either mode, image data leaves your machine unless that endpoint is local. Nothing beyond the harness's own attachment storage persists an image.
MIT
FAQs
Adaptive image routing for DeepSeek Harness: pass images directly to native multimodal models and transcribe them only for text-only models.
The npm package dsh-vision-recognizer receives a total of 559 weekly downloads. As such, dsh-vision-recognizer popularity was classified as not popular.
We found that dsh-vision-recognizer demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Product
Enterprise security teams can now detect malware, credential theft, suspicious network activity, and risky updates across Microsoft Edge extensions.

Research
/Security News
Socket researchers found 18 Chrome extensions and one Edge extension delivering a wallet drainer, credential theft, and other malicious payloads.

Product
Create ClickUp tasks from Socket alerts, automate ticketing with custom rules, and keep alert and task status synchronized.