Flyvhidbwo/dsh-vision-proxy

DeepSeek Harness plug-in: DeepSeek brain + automatic image recognition

dsh-vision-proxy closes a DSH GUI limitation: DeepSeek Harness gates image attachments on the selected model's declared inputModalities, so text-only DeepSeek models (V4-Pro, plain Flash) reject image pastes by design. Tool-based vision plugins exist, but GUI attachments still fail with text-only models selected. dsh-vision-proxy registers a new provider route (deepseek-vision) that claims image input — admitting GUI attachments — then transcribes every attached image to descriptive text in the request stream before delegating the conversation to DeepSeek. The conversation is still answered by DeepSeek; vision is an add-on transcription layer. Features: no hangs (anonymous endpoints hard-capped at 20 s; HTTP 429 fails immediately; failed endpoints cooled 60 s); multi-provider fallback chain (any OpenAI-compatible VLM endpoint, per-entry baseURL/model); zero-config local path (Ollama auto-detected at startup, no key/account needed; images never leave the machine); content-hash cache (same image transcribed at most once per process, SHA-256 keyed, max 200 entries); auto-downscale with sharp (optional; degrades gracefully without it); classified failure errors (rate_limit/quota/auth/region/model_not_found/context_too_large/http) with actionable hints; install-time consent prompt; PRIVACY NOTICE printed at startup naming the active endpoint; read_image tool compatible.

Infrastructure & Deployment ★ 11 updated 2026-08-26 ✅ runtime-tested
View on GitHub ↗

Install

dsh plugin --profile web add dsh-vision-proxy

Install from npm: dsh plugin --profile web add dsh-vision-proxy (npm 0.4.1, registry-verified 2026-08-23). During install a postinstall prompt asks whether you have a VLM API key — answer y for a paid endpoint (DashScope/Qwen, OpenRouter, etc.) or N (default) for the zero-config local path (auto-detects Ollama at localhost:11434). pnpm >= 10 blocks lifecycle build scripts by default: if the first install exits with 'Ignored build scripts: dsh-vision-proxy, sharp', approve both packages in the profile's pnpm-workspace.yaml allowBuilds and re-run. Restart dsh web, select DeepSeek + vision route in the model selector, then paste images into any conversation. Remove with dsh plugin --profile web remove dsh-vision-proxy.

Compatibility

DeepSeek Harness 0.1.0-rc.6+ (node.js >= 22.19). Registers a new provider route (deepseek-vision) that wraps the standard DeepSeek adapter and claims image input modality so GUI image attachments are admitted by DSH's preflight. As of DSH 0.1.1, DeepSeek natively supports multimodal models (e.g. DeepSeek-V4-Flash-Vision-Exp) — if you select an official vision model, no plugin is needed. dsh-vision-proxy is designed for text-only DeepSeek models (DeepSeek-V4-Pro and plain Flash), local Ollama, and custom OpenAI-compatible VLM setups. VLM backends: DashScope/Qwen (default key via VISION_API_KEY or DASHSCOPE_API_KEY), QwenCloud international, Zhipu, OpenRouter, local Ollama (auto-detected, zero config). Each fallbackModels entry can carry its own baseURL and model for chaining providers.

Details

Recent updates

The current English README documents: why the plugin exists (DSH inputModalities gate blocks GUI image attachments for text-only models), the transcription architecture, the deepseek-vision provider route, DSH 0.1.1 native multimodal note, no-hang guarantees (20 s cap, 429 fast-fail, 60 s cooldown), multi-provider fallback chain, zero-config Ollama path, content-hash cache (SHA-256, 200 entries), auto-downscale with optional sharp, classified error types with actionable hints, install-time consent prompt, PRIVACY NOTICE, pnpm >= 10 allowBuilds note, read_image compatibility.

FAQ

Do I need a vision API key?
No — with autoLocalOllama (default on), a running Ollama at localhost:11434 is detected at startup and used for free local transcription; images never leave your machine. A key is optional for paid providers (DashScope, OpenRouter, etc.).
When is this plugin NOT needed?
If you use an official DeepSeek vision model such as DeepSeek-V4-Flash-Vision-Exp (natively multimodal since DSH 0.1.1), just attach images directly — no plugin needed. This plugin is for text-only models (V4-Pro, plain Flash) and local Ollama setups.
What happens if transcription fails?
dsh-vision-proxy classifies the failure (rate_limit/quota/auth/region/model_not_found/context_too_large/http) and prints actionable hints (e.g. configure VISION_API_KEY or install Ollama) — it never silently stalls.

Alternatives

jyh20030112/dsh-visual-plugin · omdsh-dev-dsh-vision · oil-oil/dsh-vision

More plugins in Infrastructure & Deployment

Browse more in Infrastructure & Deployment

Guides for Infrastructure & Deployment plugins