54xkeee/dsh-vision
Vision for text-only DeepSeek via Doubao Web by default (zero-cost, no API key — drives your logged-in Chrome through a Windows CDP bridge), with Antigravity IDE quota (flash/pro) or Gemini fallback; auto detail escalation, vision evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.
dsh-vision gives text-only DeepSeek models on DeepSeek Harness the ability to see images — powered by Doubao Web by default at zero cost with no API key (log into Doubao once in a browser driven by the bundled Windows bridge). When you paste an image, dsh-vision turns it into a text placeholder; DeepSeek calls the vision tool; the chosen channel recognizes it; text evidence flows back and the answer continues. The vision engine is a structured, self-escalating, memory-backed pipeline: auto detail escalation (a standard pass first, auto-upgrade to a deep pass for complex scenes like OCR-heavy content, tables, charts, UI screens, counting, or comparison), four task modes (glance, ocr, region with normalized coords or plain language, compare between ≥2 images with confidence), strict JSON evidence objects that separate observations from inference and mark uncertainty explicitly, long-context visual memory (every result written to the session timeline as a durable record, reused across turns, restored after compaction), and content-hash caching so the same image + same question is recognized once per process. Alternate channels cover Antigravity IDE quota, Gemini API, Cockpit proxy, and any IDE CLI (Claude Code, Gemini CLI, Qwen Code, MiMo).
Install
dsh plugin --profile web add dsh-vision-webnpm package dsh-vision-web 0.1.0 (registry-verified 2026-08-24; README badge matches). Install: dsh plugin --profile web add dsh-vision-web. Default channel is Doubao Web — zero cost, no API key: log into doubao.com once in a Chrome driven by the included Windows bridge (node bridge.mjs, puppeteer-core + CDP), restart dsh web, and paste images. Other channels: Antigravity IDE quota (flash/pro tiers), Gemini API (genlangKey), Cockpit proxy, any IDE CLI (Claude Code / Gemini CLI / Qwen Code / MiMo via ideCli config). Requires Node.js ≥ 20. The plugin wraps text-only DeepSeek models so pasted images become text placeholders; the model calls the vision tool, evidence flows back, and the answer continues. Includes vision evidence memory (results persist in the session, reused across turns, restored after compaction) and content-hash caching (same image + same question = recognized once per process).
Compatibility
DeepSeek Harness web profile, Node.js ≥ 20. Vision channels: Doubao Web (default, zero cost, browser-login based, Windows bridge with WSL queue at 127.0.0.1:9340), Antigravity IDE quota (flash/pro tiers by model name, auto-discovered ports/CSRF), Gemini API, Cockpit proxy, aicode direct channel (bring your own OAuth), and any IDE CLI. Auto fallback chain with winCurl fallback for WSL/firewalled networks. Detail escalation (auto: standard pass first, deep pass for complex scenes), four task modes (glance/ocr/region/compare), strict JSON evidence objects, long-context visual memory with compaction rehydration, content-hash LRU cache (default 64). Privacy: Doubao Web images go to your own logged-in Doubao session; API keys stay in local config with redacted error messages.
Details
- Repo: 54xkeee/dsh-vision
- Category: Vision & Multimodal
- Stars: 7
- Version: npm package dsh-vision-web 0.1.0 (registry-verified 2026-08-24)
- Last push: 2026-08-16
- First seen: 2026-08-16
Recent updates
The current bilingual README (English edition) documents: the zero-cost Doubao Web default, quick start paths (Doubao Web / Antigravity / Gemini / any IDE CLI), the Doubao bridge architecture (WSL queue + Windows puppeteer-core bridge), the vision engine (auto detail escalation, four task modes, structured evidence, long-context memory, content-hash caching), full configuration reference, troubleshooting, privacy, architecture, and development.
FAQ
- Does it require an API key?
- Not for the default channel — Doubao Web automates your logged-in browser, so it is zero cost with no API key: log into doubao.com once in the bridge's Chrome profile and the bridge reuses it forever. Paid/API channels (Antigravity, Gemini, aicode) are optional alternatives.
- How does the model remember what it saw?
- Every recognition result is written to the session timeline as a durable dsh-vision-evidence record. The same image + same question hits the existing record, and after DSH compacts a long session, recent vision records are restored automatically.
- How does it handle complex images like screenshots or tables?
- detail: auto runs a two-pass strategy — a standard pass with triage, and if the vision model classifies the scene as complex (dense small text, tables/charts/code, UI screens, counting, comparison, multi-subject relationships), a deep pass is triggered automatically.
Alternatives
Flyvhidbwo/dsh-vision-proxy · oil-oil/dsh-vision · Anionex/dsh-vision-toolkit