54xkeee/dsh-vision

Vision for text-only DeepSeek via Doubao Web by default (zero-cost, no API key — drives your logged-in Chrome through a Windows CDP bridge), with Antigravity IDE quota (flash/pro) or Gemini fallback; auto detail escalation, vision evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.

dsh-vision gives text-only DeepSeek models on DeepSeek Harness the ability to see images — powered by Doubao Web by default at zero cost with no API key (log into Doubao once in a browser driven by the bundled Windows bridge). When you paste an image, dsh-vision turns it into a text placeholder; DeepSeek calls the vision tool; the chosen channel recognizes it; text evidence flows back and the answer continues. The vision engine is a structured, self-escalating, memory-backed pipeline: auto detail escalation (a standard pass first, auto-upgrade to a deep pass for complex scenes like OCR-heavy content, tables, charts, UI screens, counting, or comparison), four task modes (glance, ocr, region with normalized coords or plain language, compare between ≥2 images with confidence), strict JSON evidence objects that separate observations from inference and mark uncertainty explicitly, long-context visual memory (every result written to the session timeline as a durable record, reused across turns, restored after compaction), and content-hash caching so the same image + same question is recognized once per process. Alternate channels cover Antigravity IDE quota, Gemini API, Cockpit proxy, and any IDE CLI (Claude Code, Gemini CLI, Qwen Code, MiMo).

Vision & Multimodal ★ 7 updated 2026-08-16
View on GitHub ↗

Install

dsh plugin --profile web add dsh-vision-web

npm package dsh-vision-web 0.1.0 (registry-verified 2026-08-24; README badge matches). Install: dsh plugin --profile web add dsh-vision-web. Default channel is Doubao Web — zero cost, no API key: log into doubao.com once in a Chrome driven by the included Windows bridge (node bridge.mjs, puppeteer-core + CDP), restart dsh web, and paste images. Other channels: Antigravity IDE quota (flash/pro tiers), Gemini API (genlangKey), Cockpit proxy, any IDE CLI (Claude Code / Gemini CLI / Qwen Code / MiMo via ideCli config). Requires Node.js ≥ 20. The plugin wraps text-only DeepSeek models so pasted images become text placeholders; the model calls the vision tool, evidence flows back, and the answer continues. Includes vision evidence memory (results persist in the session, reused across turns, restored after compaction) and content-hash caching (same image + same question = recognized once per process).

Compatibility

DeepSeek Harness web profile, Node.js ≥ 20. Vision channels: Doubao Web (default, zero cost, browser-login based, Windows bridge with WSL queue at 127.0.0.1:9340), Antigravity IDE quota (flash/pro tiers by model name, auto-discovered ports/CSRF), Gemini API, Cockpit proxy, aicode direct channel (bring your own OAuth), and any IDE CLI. Auto fallback chain with winCurl fallback for WSL/firewalled networks. Detail escalation (auto: standard pass first, deep pass for complex scenes), four task modes (glance/ocr/region/compare), strict JSON evidence objects, long-context visual memory with compaction rehydration, content-hash LRU cache (default 64). Privacy: Doubao Web images go to your own logged-in Doubao session; API keys stay in local config with redacted error messages.

Details

Recent updates

The current bilingual README (English edition) documents: the zero-cost Doubao Web default, quick start paths (Doubao Web / Antigravity / Gemini / any IDE CLI), the Doubao bridge architecture (WSL queue + Windows puppeteer-core bridge), the vision engine (auto detail escalation, four task modes, structured evidence, long-context memory, content-hash caching), full configuration reference, troubleshooting, privacy, architecture, and development.

FAQ

Does it require an API key?
Not for the default channel — Doubao Web automates your logged-in browser, so it is zero cost with no API key: log into doubao.com once in the bridge's Chrome profile and the bridge reuses it forever. Paid/API channels (Antigravity, Gemini, aicode) are optional alternatives.
How does the model remember what it saw?
Every recognition result is written to the session timeline as a durable dsh-vision-evidence record. The same image + same question hits the existing record, and after DSH compacts a long session, recent vision records are restored automatically.
How does it handle complex images like screenshots or tables?
detail: auto runs a two-pass strategy — a standard pass with triage, and if the vision model classifies the scene as complex (dense small text, tables/charts/code, UI screens, counting, comparison, multi-subject relationships), a deep pass is triggered automatically.

Alternatives

Flyvhidbwo/dsh-vision-proxy · oil-oil/dsh-vision · Anionex/dsh-vision-toolkit

More plugins in Vision & Multimodal

Browse more in Vision & Multimodal

Guides for Vision & Multimodal plugins