Yuuz12/dsh-vision-helper
DeepSeek Harness Vision Helper/DeepSeek Harness Vision Assistance Solution
A persistent vision plugin registering the vision_analyze tool: the agent analyzes images (OCR, UI/error screenshots, chart understanding, general image description) through any configurable multimodal model, with a settings-page UI. Choose provider/model in Settings → 视觉助手 (or edit dsh-vision-helper.json in the data dir); leaving it empty auto-selects a multimodal model. Robustness: strips invisible Unicode chars from pasted paths (U+202A etc.), caps image size by longest edge (maxEdge, default 4096px), dedupes streaming text, friendly errors. A system-prompt segment nudges the agent to call vision_analyze for image-related tasks. Persistent and global — survives restarts, all sessions.
Install
npx @deepseek-ai/dsh plugin --profile web add dsh-vision-helpernpm install (recommended): npx @deepseek-ai/dsh plugin --profile web add dsh-vision-helper (npm 0.4.2 registry-verified 2026-08-29) — pure JS package, no prepare script, no build authorization needed, bundle layer auto-activated. GitHub: npx @deepseek-ai/dsh plugin --profile web add github:Yuuz12/dsh-vision-helper. Manual: copy the folder into <profile>/node_modules/ + add the plugin row to cordis.patch.yml. Verify with npx @deepseek-ai/dsh --profile web --dump-config and restart. English README: README.en.md.
Compatibility
DSH Web deployment (settings page is Web-only; the tool works in any profile mounting the plugin). Zero-dependency host module; requires a multimodal model route configured in llm-pi-ai settings (or any OpenAI-compatible vision endpoint). No API key bundled.
Details
- Repo: Yuuz12/dsh-vision-helper
- Category: Other
- Stars: 2
- Version: npm 0.4.2 (registry-verified 2026-08-29)
- Last push: 2026-08-15
- First seen: 2026-08-13
Recent updates
vision_analyze tool via any configurable multimodal model; settings page; auto multimodal-model selection; invisible-Unicode stripping; 4096px max edge.
FAQ
- What can the agent do with it?
- Call vision_analyze with a local image path or data: URI to get analysis text — OCR, screenshots, charts, or general description.
- Do I need an API key from this plugin?
- No — configure a multimodal model route in llm-pi-ai settings (or any OpenAI-compatible vision endpoint); the plugin carries no key.
- Is it persistent?
- Yes — a host composition plugin: survives restarts and is available in all sessions.
Alternatives
Scorp1o117/dsh-tool-vision · RRRosmontis/dsh-qwen-mm · libinyam/dsh-vision-provider