TZHR-invest/dsh-plugins#dsh-vision-tool

Agent-callable vision tool that describes local images via any OpenAI-compatible vision endpoint you configure, with an optional multi-model cross-check and no built-in keys.

An agent-callable vision tool for the DeepSeek Harness: it describes local images on demand via any OpenAI-compatible multimodal model, so text-only models can 'see' files in the workspace by handing their paths to this tool.

Vision & Multimodal ★ 2 updated 2026-08-17
View on GitHub ↗

Install

dsh plugin --profile web add dsh-vision-tool

dsh plugin --profile web add dsh-vision-tool (npm dsh-vision-tool 0.1.1, verified npm 2026-08-31; part of the TZHR-invest/dsh-plugins monorepo).

Compatibility

DSH web profile; uses any OpenAI-compatible multimodal model endpoint configured by the user.

Details

Recent updates

Agent-callable vision tool; describes local images via OpenAI-compatible multimodal endpoint; path-based workflow.

FAQ

Which model does it use?
Any OpenAI-compatible multimodal model you configure — the endpoint and key are user-supplied, so it works with qwen-vl, gpt-4o, GLM-4V, and similar providers.
Can text-only models use it?
Yes — that is the point: the tool is agent-callable, so a text-only model can invoke it with a local image path and receive a description.
Does it upload images anywhere?
Images are sent to the configured multimodal endpoint for analysis; no other egress is involved.

Alternatives

RRRosmontis/dsh-qwen-mm · Yuuz12/dsh-vision-helper · 54xkeee/dsh-youreyes

More plugins in Vision & Multimodal

Browse more in Vision & Multimodal

Guides for Vision & Multimodal plugins