imkingjh999/dsh-tool-accurate-vision
Model-facing accurate_vision tool for DeepSeek Harness: precise spatial reasoning via any OpenAI-compatible vision model (0-1000 bbox primitives + annotated SVG)
dsh-tool-accurate-vision adds a model-facing accurate_vision tool (ported from pi-accurate-vision) for precise spatial reasoning over an image: a vision model reads the image and returns a structured note plus bounding-box primitives normalized to 0–1000, which the tool formats as a <vision-context> block the next model turn reads — giving a text-only agent exact object positions, layout and OCR without losing spatial fidelity. Every call also writes a self-contained SVG with every bounding box and label drawn on the original image, returned as an annotatedImage path so boxes can be eyeballed instead of trusted blind (set annotate: false to skip). The pure vision core is provider-agnostic; the Cordis host owns config, credential resolution and the registered tool.
Install
dsh plugin --profile web add dsh-tool-accurate-visionnpm dsh-tool-accurate-vision 0.1.0 verified 2026-09-03 (repository field → github.com/imkingjh999/dsh-tool-accurate-vision; README bilingual EN/中文). Install: dsh plugin --profile web add dsh-tool-accurate-vision. Requires a vision API key separate from DEEPSEEK_API_KEY: export VISION_API_KEY=sk-... (any OpenAI-compatible multimodal endpoint works).
Compatibility
DSH with text-only main models; any OpenAI-compatible multimodal chat/completions endpoint as the vision backend; VISION_API_KEY env or config.
Details
- Repo: imkingjh999/dsh-tool-accurate-vision
- Category: uncategorized
- Stars: 0
- Version: npm dsh-tool-accurate-vision 0.1.0
- Last push: 2026-08-17
- First seen: 2026-08-17
Recent updates
0.1.0 current on npm (verified 2026-09-03).
FAQ
- Which vision backends work?
- Any OpenAI-compatible multimodal chat/completions endpoint — the core bridge is provider-agnostic and takes the API key via VISION_API_KEY.
- How does a text-only model get spatial info?
- The vision model returns bounding-box primitives normalized to 0–1000; the tool wraps them into a <vision-context> block for the next turn.
- Can I verify what it 'saw'?
- Yes — each call writes an SVG with every box and label drawn on the original image (annotatedImage), unless you set annotate: false.
Alternatives
good-boy4069/dsh-vision-guard · 1710782766/dsh-llm-vision