Scorp1o117/dsh-tool-vision
Vision model for DeepSeek Harness | DeepSeek Harness external vision model plug-in
dsh-tool-vision gives DeepSeek Harness a vision model through any OpenAI-compatible API: configure baseURL, apiKey/apiKeyEnv, and model (gpt-4o-mini by default), and the model gains image understanding for screenshots, errors, UI analysis, OCR, and documents. bridgeTextOnly mode bridges pasted images to text hints on models that cannot see images, exporting bridged images to a temp dir, while multimodalModels lists model ids that receive image blocks directly. It is part of the DeepSeek Harness Enhancement Suite (Vision · Soul/Persona · Long-term Memory · Plugin Marketplace).
Install
dsh plugin --profile web add dsh-tool-visionnpm package dsh-tool-vision 0.6.4 (registry-verified 2026-08-24). Install: dsh plugin --profile web add dsh-tool-vision, then mount it in the profile patch ($DSH_HOME/profiles/<name>/cordis.patch.yml) with an insert row (id: tool-vision, name: 'dsh-tool-vision', config: baseURL, apiKeyEnv, model). Or load from a local path without npm (name: './plugins/dsh-tool-vision/index.js'). Part of the DeepSeek Harness Enhancement Suite. Config defaults: baseURL https://api.openai.com/v1, apiKeyEnv VISION_API_KEY, model gpt-4o-mini, maxTokens 1024, timeoutMs 60000, maxImageBytes 10MB, bridgeTextOnly true, bridgeExportDir os.tmpdir()/dsh-vision-bridge.
Compatibility
DeepSeek Harness web profile. OpenAI-compatible vision API: baseURL (default https://api.openai.com/v1), apiKey (takes precedence over env) or apiKeyEnv (default VISION_API_KEY), model (default gpt-4o-mini), maxTokens (1024), timeoutMs (60000), maxImageBytes (10MB). bridgeTextOnly (default true) bridges pasted images to text hints on models that cannot see images, exporting bridged images to bridgeExportDir (os.tmpdir()/dsh-vision-bridge); multimodalModels lists model ids that receive image blocks directly. Part of the Enhancement Suite.
Details
- Repo: Scorp1o117/dsh-tool-vision
- Category: Other
- Stars: 6
- Version: npm package dsh-tool-vision 0.6.4 (registry-verified 2026-08-24)
- Last push: 2026-09-11
- First seen: 2026-08-13
Recent updates
The current README documents: install (dsh plugin add + profile patch mount, or local-path load), the full config table (baseURL, apiKey, apiKeyEnv, model, maxTokens, timeoutMs, maxImageBytes, description, bridgeTextOnly, bridgeExportDir, multimodalModels), and the Enhancement Suite relationship.
FAQ
- Which API does it use?
- Any OpenAI-compatible vision API — configure baseURL (default https://api.openai.com/v1), apiKey or apiKeyEnv (default VISION_API_KEY), and model (default gpt-4o-mini).
- How does it work with text-only models?
- bridgeTextOnly (default true) bridges pasted images to text hints on models that cannot see images, exporting bridged images to a temp dir; multimodalModels lists model ids that receive image blocks directly.
- Do I need npm to load it?
- No — it can also be loaded from a local path without npm (name: './plugins/dsh-tool-vision/index.js').
Alternatives
54xkeee/dsh-vision · FuzzySoul/dsh-free-vision · xsoc1/dsh-image-vision