gloryxpnv/dsh-tool-vision

Local-first structured vision for text-only agents: images go to a local OpenAI-compatible VLM and come back as JSON evidence (summary, verbatim OCR, layout regions, entities/relations, colors, explicit uncertainty), with anti-hallucination fallback and an optional paste/upload bridge; zero cloud cost, images never leave the machine.

dsh-tool-vision gives a text-only model (DeepSeek, GLM, or any chat model without image input) the ability to see, using a local vision model. The VLM is asked to fill a fixed JSON evidence template — summary, verbatim OCR text and lines, layout regions with reading order, semantics (entities and relations), visual colors/style, and an explicit uncertainty list — so the main model quotes specifics instead of guessing. Anti-hallucination by design: the template requires the model to state what it could not determine, OCR of a textless image returns an empty field rather than invented words, and an invalid-JSON reply falls back to the raw answer marked as such. One plugin covers two surfaces: a model-facing vision tool plus an optional vision-bridge service that lets text-only routes admit pasted/uploaded images (with keepThumbnail and on-demand autoDescribe options).

Vision & Multimodal ★ 0 updated 2026-08-15
View on GitHub ↗

Install

dsh plugin --profile web add dsh-vision-local

The README's install line is 'dsh plugin --profile web add dsh-vision-local' run in a DSH profile directory (or via the dsh CLI); you then add the plugin row to the profile patch (cordis.patch.yml) or rely on the bundle's own layer. The npm package dsh-vision-local 0.3.0 was re-verified live on 2026-10-02; its registry entry carries NO repository field, but the README states the npm package name is dsh-vision-local and the source repo is gloryxpnv/dsh-tool-vision.

Compatibility

DeepSeek Harness, web and headless/text-only routes. Works with any local OpenAI-compatible VLM endpoint (LM Studio, Ollama, vLLM); default endpoint is http://127.0.0.1:1234/v1 (LM Studio's default port). Defaults are tuned for a local workstation GPU running a 9B-class VLM (8192 output tokens, 50 MB image cap, 180 s timeout). Zero API cost and no image bytes leave the machine.

Details

Recent updates

The README documents the local-only architecture diagram, the structured-evidence JSON template (summary / ocr / layout / semantics / visual / uncertainty), the anti-hallucination fallback, the vision-bridge paste/upload path with keepThumbnail + autoDescribe, the zero-config defaults, and the config fields (endpoint, model id, token budget, timeouts, image size cap, structured on/off).

FAQ

How do I install dsh-tool-vision?
dsh plugin --profile web add dsh-vision-local (npm dsh-vision-local 0.3.0 re-verified 2026-10-02), then add the plugin row to cordis.patch.yml or rely on the bundle's own layer.
Does it send my images to the cloud?
No — the README says images go only to your own local vision model (LM Studio / Ollama / any OpenAI-compatible endpoint); no cloud keys and no image bytes leave the machine.
What happens if the local VLM returns invalid JSON?
The plugin falls back to the raw answer and marks it — it never silently fabricates; the template also requires an explicit uncertainty list.

Alternatives

siegfly/dsh-deepseek-vision · jiavenzhong/dsh-tool-markitdown · welsione/dsh-mmx-bridge

More plugins in Vision & Multimodal

Browse more in Vision & Multimodal

Guides for Vision & Multimodal plugins