whitelonng/dsh-plugin-describe-image
DeepSeek Harness plugin: describe_image — give a text-only model vision through an OpenAI-compatible VLM endpo
dsh-plugin-describe-image gives text-only models (DeepSeek V4 and friends) the ability to understand images: the agent or user hands the describe_image tool an image, the tool asks a configured vision-language model at an OpenAI-compatible endpoint to describe it, and only the description text goes back into the conversation and session log. Security and bounds: redirects refused on every request, maxBytes / maxOutputTokens / timeoutMs bounds, magic-byte media-type gate, bounded error excerpts, secrets never logged.
Install
dsh plugin --profile web add github:whitelonng/dsh-plugin-describe-imageGitHub source install per README (no npm package documented): dsh plugin --profile web add github:whitelonng/dsh-plugin-describe-image. The desktop app's plugin list accepts the same spec in its install box; the plugin loads after an application restart. This repo holds the plugin subtree as it lives inside deepseek-harness (packages/vision/tool-describe-image) — the harness tree is the build environment. Companion harness changes (image-flattening on text-only routes) ship in the harness repo, not this subtree.
Compatibility
DeepSeek Harness. Adds the model-facing describe_image tool: loads one image (local path, http(s) URL, or durable attachment reference) and asks a vision-language model at an OpenAI-compatible endpoint (Qwen-VL, GLM-4V, GPT-4o, or local Ollama) to describe it; only the returned text crosses into the conversation — the image never enters the session log. Live config card under Settings → Plugins → 'Image understanding' edits baseURL, model, and API key with immediate effect. Per-call key resolution: inline apiKey → credential seam (apiKeyEnv, default VISION_API_KEY) → launch environment.
Details
- Repo: whitelonng/dsh-plugin-describe-image
- Category: Other
- Stars: 6
- Version: GitHub source install github:whitelonng/dsh-plugin-describe-image (README-documented)
- Last push: 2026-08-15
- First seen: 2026-08-14
Recent updates
The current English README documents: install (dsh plugin add github:...), features (three input forms, live configuration card, per-call API key resolution, security and bounds), a quick-start cordis.yml example, FAQ (which vision models work, whether the image enters the conversation, API key configuration, malicious-input safety), the repository layout, and the git subtree sync workflow.
FAQ
- Does the image itself enter the conversation?
- No. The image is loaded, checked, and sent only to the vision endpoint; the session log and the model see only the returned description text.
- Which vision models work?
- Any OpenAI-compatible vision endpoint: Qwen-VL (dashscope.aliyuncs.com/compatible-mode/v1), GLM-4V, GPT-4o, or a local Ollama endpoint. Set baseURL and model in Settings → Plugins → 'Image understanding'.
- How is the API key resolved?
- Three layers in order: an inline apiKey in config, the credential seam (apiKeyEnv, default VISION_API_KEY), then the launch environment. The key is never written into logs.
Alternatives
54xkeee/dsh-vision · oil-oil/dsh-vision · Anionex/dsh-vision-toolkit