MJorgin/dsh-media-skills

dsh-media-skills gives DeepSeek Harness free image reading and generation: paste/drag/pick an image in any text-only session and a free vision model turns it into text your current model understands (Powered by the Harness core auto-description path), plus a vision-review skill that analyzes images/screenshots (catch UI visual bugs, detect watermarks, turn images into text, optional ModLens-style structured evidence output) and a media-tools skill that generates illustrations, avatars, backgrounds and banners with a free watermark-free model (SenseNova U1 Fast → SiliconFlow Kolors). Engine failover: GLM-4V-Flash (free) → DeepSeek-V4-Flash-Vision-Exp (same key, higher quality) → SiliconFlow Qwen3-VL → SenseNova → Google Gemini → any OpenAI-compatible endpoint. No hardcoded keys, no paid API, no file saving, no session switching; keys are never stored in the repo (env vars → ~/.dsh/secrets/media-tools.env → ~/.codex/secrets/media-tools.env).

Agent Capabilities ★ 14 updated 2026-09-21 ✅ runtime-tested
View on GitHub ↗

Install

dsh plugin --profile web add github:MJorgin/dsh-media-skills

dsh plugin --profile <name> add github:MJorgin/dsh-media-skills. Keys: on v0.1.1-rc.1+ zero extra keys — paste reading and the vision route run on your agent's existing DEEPSEEK_API_KEY (DeepSeek-V4-Flash-Vision-Exp). On rc.7/rc.8 (or to add free engines) add Zhipu (open.bigmodel.cn → API Keys, glm-4v-flash is free), SiliconFlow (siliconflow.cn → API Keys, Kolors is free), optional Google Gemini (aistudio.google.com) — via the Web GUI Settings → Models or the credentials file (~/.dsh/.credentials.yaml with GLM_API_KEY). Restart dsh web and hard-refresh; verify the model selector shows 智谱 GLM-4V-Flash(视觉).

Compatibility

DeepSeek Harness with skill support; Python 3.9+; vision model route GLM-4V-Flash (free) + DeepSeek-V4-Flash-Vision-Exp (same key as agent on v0.1.1) with failover chain to SiliconFlow Qwen3-VL → SenseNova → Google Gemini → any OpenAI-compatible endpoint. Paste-image reading requires a Harness build with core api-proxy image-admission support (bundled patches for rc.7/rc.8/v0.1.1-rc.1).

Details

Recent updates

The current README documents the two skills (vision-review, media-tools), the engine failover chain, the zero-extra-key v0.1.1 path vs rc.7/rc.8 patched path, key storage locations, the ModLens comparison/coexistence, and the 9-language docs.

FAQ

Does paste-image reading need a Harness core patch?
The auto-describe pipeline lives in Harness core (api-proxy image-admission logic); the bundle ships the model route + skills. The vision model works on any DSH build, but paste reading requires a build with that core support (patches bundled for rc.7/rc.8/v0.1.1-rc.1).
Is media-tools really free?
Yes — SiliconFlow Kolors is free and watermark-free; if a model is temporarily disabled the skill lists available models and you can switch.
Where are API keys stored?
Never in the repo: skill scripts read env vars → ~/.dsh/secrets/media-tools.env → ~/.codex/secrets/media-tools.env, and the vision route reads GLM_API_KEY from DSH's credential store.

Alternatives

william-jin-cmu/dsh-vision · oil-oil/dsh-vision · liustack/modlens

More plugins in Agent Capabilities

Browse more in Agent Capabilities

Guides for Agent Capabilities plugins