whitelonng/dsh-plugin-describe-image

DeepSeek Harness plugin: describe_image — give a text-only model vision through an OpenAI-compatible VLM endpo

dsh-plugin-describe-image gives text-only models (DeepSeek V4 and friends) the ability to understand images: the agent or user hands the describe_image tool an image, the tool asks a configured vision-language model at an OpenAI-compatible endpoint to describe it, and only the description text goes back into the conversation and session log. Security and bounds: redirects refused on every request, maxBytes / maxOutputTokens / timeoutMs bounds, magic-byte media-type gate, bounded error excerpts, secrets never logged.

Other ★ 6 updated 2026-08-15 ⚠️ needs adapt
View on GitHub ↗

Install

dsh plugin --profile web add github:whitelonng/dsh-plugin-describe-image

GitHub source install per README (no npm package documented): dsh plugin --profile web add github:whitelonng/dsh-plugin-describe-image. The desktop app's plugin list accepts the same spec in its install box; the plugin loads after an application restart. This repo holds the plugin subtree as it lives inside deepseek-harness (packages/vision/tool-describe-image) — the harness tree is the build environment. Companion harness changes (image-flattening on text-only routes) ship in the harness repo, not this subtree.

Compatibility

DeepSeek Harness. Adds the model-facing describe_image tool: loads one image (local path, http(s) URL, or durable attachment reference) and asks a vision-language model at an OpenAI-compatible endpoint (Qwen-VL, GLM-4V, GPT-4o, or local Ollama) to describe it; only the returned text crosses into the conversation — the image never enters the session log. Live config card under Settings → Plugins → 'Image understanding' edits baseURL, model, and API key with immediate effect. Per-call key resolution: inline apiKey → credential seam (apiKeyEnv, default VISION_API_KEY) → launch environment.

Details

Recent updates

The current English README documents: install (dsh plugin add github:...), features (three input forms, live configuration card, per-call API key resolution, security and bounds), a quick-start cordis.yml example, FAQ (which vision models work, whether the image enters the conversation, API key configuration, malicious-input safety), the repository layout, and the git subtree sync workflow.

FAQ

Does the image itself enter the conversation?
No. The image is loaded, checked, and sent only to the vision endpoint; the session log and the model see only the returned description text.
Which vision models work?
Any OpenAI-compatible vision endpoint: Qwen-VL (dashscope.aliyuncs.com/compatible-mode/v1), GLM-4V, GPT-4o, or a local Ollama endpoint. Set baseURL and model in Settings → Plugins → 'Image understanding'.
How is the API key resolved?
Three layers in order: an inline apiKey in config, the credential seam (apiKeyEnv, default VISION_API_KEY), then the launch environment. The key is never written into logs.

Alternatives

54xkeee/dsh-vision · oil-oil/dsh-vision · Anionex/dsh-vision-toolkit

More plugins in Other

Browse more in Other

Guides for Other plugins