FuzzySoul/dsh-free-vision
Free vision bridge for text-only models: image understanding, OCR, UI and debug analysis via free-tier providers (Qwen3-VL-Flash, Doubao, DeepSeek-OCR) with a settings GUI.
dsh-free-vision gives text-only models on DeepSeek Harness image understanding using free-tier vision models with zero MCP configuration: paste a screenshot, error, UI, document, or photo and the image_understand tool (a single general tool registered on ctx.tools) recognizes it via the chosen provider — Qwen3-VL-Flash by default at zero cost. It supports task modes (auto/general/ocr/ui/debug/describe), auto detail escalation with multi-cropping for large images, per-provider API base URL overrides (proxy, gateway, local service, or any OpenAI-compatible endpoint), and direct connections that strip proxy variables for domestic API access. A 1 MB screenshot is roughly 2,600 tokens, and the qwen free quota covers about 190,000 images.
Install
dsh plugin --profile web add dsh-free-visionnpm package dsh-free-vision 1.0.8 (registry-verified 2026-08-24). Install: dsh plugin --profile web add dsh-free-vision. Restart dsh web — the image_understand tool becomes available (rename via config.toolName). Zero MCP configuration: the vision engine (luma-mcp) is bundled as a dependency and started in-process. After restart, open Settings → Free Vision for the config form (API keys, provider, tool name) — settings are saved to ~/.dsh/free-vision.json and take effect immediately on the next call without a restart.
Compatibility
DeepSeek Harness web profile. Free-first vision providers: qwen (default, Qwen3-VL-Flash via DASHSCOPE_API_KEY — Alibaba Cloud free tier ~500k tokens), volcengine (Doubao vision, VOLCENGINE_API_KEY, free 200k+ tokens), siliconflow (DeepSeek-OCR free, SILICONFLOW_API_KEY), plus zhipu (GLM-4.6V), hunyuan (HY-Vision), and custom OpenAI-compatible endpoints (CUSTOM_API_KEY + CUSTOM_BASE_URL + CUSTOM_MODEL_NAME). Every provider can override its API base URL (proxy/gateway/local service). Direct connection: the subprocess strips proxy environment variables for direct domestic-API access. Task modes: auto | general | ocr | ui | debug | describe; large images are auto multi-cropped for fidelity. Bilingual tool descriptions (EN/中文).
Details
- Repo: FuzzySoul/dsh-free-vision
- Category: Vision & Multimodal
- Stars: 6
- Version: npm package dsh-free-vision 1.0.8 (registry-verified 2026-08-24)
- Last push: 2026-08-19
- First seen: 2026-08-15
Recent updates
The current bilingual README (English edition) documents: why free (provider free-tier table with quota estimates), features (zero MCP config, single image_understand tool, free-first multi-provider, per-provider base URL override, direct connection, task modes, bilingual), install (dsh plugin add + restart), settings UI (Settings → Free Vision, ~/.dsh/free-vision.json, immediate effect), and provider environment variables.
FAQ
- Do I need an API key?
- Yes, but the default providers are free-tier: qwen (Qwen3-VL-Flash, ~500k free tokens via DASHSCOPE_API_KEY), volcengine Doubao (free 200k+), and siliconflow DeepSeek-OCR (free).
- Do I need MCP configuration?
- No — the vision engine (luma-mcp) is bundled as a package dependency and started in-process; there is no cordis.patch.yml editing or runtime npx.
- How does it handle large or complex images?
- Task mode auto routes to general/ocr/ui/debug/describe as needed, and large images are automatically multi-cropped to preserve detail.
Alternatives
54xkeee/dsh-vision · Flyvhidbwo/dsh-vision-proxy · stardustlc666/dsh-codex-port