1710782766/dsh-llm-vision
Reliable vision + OCR for text-only models on DeepSeek Harness: describe_image (normal/critical) + extract_text tools, auto-preprocessing, retries, and a persistent answer cache.
dsh-llm-vision gives text-only DeepSeek Harness models reliable image understanding and OCR, configured entirely in the GUI. It registers model-facing tools backed by any OpenAI-compatible vision endpoint: describe_image (natural description plus a 'critical' objective-inspection perspective that reports text misalignment, overlap, occlusion and wrapping anomalies and separates fact from guess — for UI bug reports and screenshot-vs-design comparisons; accepts one image or a batch of up to 8), extract_text (OCR through a dedicated OCR model with structured JSON/CSV output and verbatim extraction that never guesses missing text), and llm_vision_check (diagnostics that verify config, key resolution and endpoint auth — the key itself never appears in the report). The browser half rewrites paste/drag/drop image sends into attach references the text model can resolve, upgrades them to inline thumbnails in the transcript, and a live settings card (Settings → Plugins → llm-vision) controls endpoint, models, prompts, bounds, retries, preprocessing and cache. Inputs may be local absolute paths, http(s) URLs (redirects refused) or attachment references; the image never enters the session log — only the returned text crosses into the conversation.
Install
dsh plugin --profile web add dsh-llm-vision@0.3.2npm dsh-llm-vision 0.3.2 verified 2026-09-03 (repository field → github.com/1710782766/dsh-llm-vision; README EN primary with zh edition). The README pins the version on purpose: pnpm 11 holds back packages published in the last 24h, so a bare add dsh-llm-vision (latest) could silently install the previous release on launch day. Restart the GUI once after install — plugins load at boot, so the plugin and its settings card appear only after restart; later config changes never need one. Requires dsh >= 0.1.2-alpha.1 (the session log format it reads).
Compatibility
dsh >= 0.1.2-alpha.1; Node per dsh; OpenAI-compatible vision endpoints (free presets for Zhipu / Gemini / DashScope; any compatible endpoint configurable in the settings card).
Details
- Repo: 1710782766/dsh-llm-vision
- Category: uncategorized
- Stars: 0
- Version: npm dsh-llm-vision 0.3.2
- Last push: 2026-08-17
- First seen: 2026-08-17
Recent updates
describe_image (normal + critical lenses, batch up to 8) / extract_text OCR / llm_vision_check diagnostics; GUI-only config; paste/drag/drop image rewrite; image bytes never enter the session log.
FAQ
- How do I start?
- Install the pinned release, restart the GUI, open Settings → Plugins → llm-vision, pick a Provider preset (zhipu / gemini / dashscope fill the endpoints — free routes), paste an API key (stored in the owner-only settings document, never shown again), then paste or drop an image into the composer.
- Which models can use it?
- Text-only models such as DeepSeek V4 or GLM text series — the plugin forwards images to a separate OpenAI-compatible vision/OCR endpoint and returns text.
- Does the image enter the session log?
- No. The image itself never crosses into the conversation; only the returned text description is recorded.
Alternatives
apheli0os/deepseek-harness-orchestrate · MoneShadow/dsh-plugin-vision · wdwind/dsh-vision-no-vision