MJorgin/dsh-media-skills
dsh-media-skills gives DeepSeek Harness free image reading and generation: paste/drag/pick an image in any text-only session and a free vision model turns it into text your current model understands (Powered by the Harness core auto-description path), plus a vision-review skill that analyzes images/screenshots (catch UI visual bugs, detect watermarks, turn images into text, optional ModLens-style structured evidence output) and a media-tools skill that generates illustrations, avatars, backgrounds and banners with a free watermark-free model (SenseNova U1 Fast → SiliconFlow Kolors). Engine failover: GLM-4V-Flash (free) → DeepSeek-V4-Flash-Vision-Exp (same key, higher quality) → SiliconFlow Qwen3-VL → SenseNova → Google Gemini → any OpenAI-compatible endpoint. No hardcoded keys, no paid API, no file saving, no session switching; keys are never stored in the repo (env vars → ~/.dsh/secrets/media-tools.env → ~/.codex/secrets/media-tools.env).
Install
dsh plugin --profile web add github:MJorgin/dsh-media-skillsdsh plugin --profile <name> add github:MJorgin/dsh-media-skills. Keys: on v0.1.1-rc.1+ zero extra keys — paste reading and the vision route run on your agent's existing DEEPSEEK_API_KEY (DeepSeek-V4-Flash-Vision-Exp). On rc.7/rc.8 (or to add free engines) add Zhipu (open.bigmodel.cn → API Keys, glm-4v-flash is free), SiliconFlow (siliconflow.cn → API Keys, Kolors is free), optional Google Gemini (aistudio.google.com) — via the Web GUI Settings → Models or the credentials file (~/.dsh/.credentials.yaml with GLM_API_KEY). Restart dsh web and hard-refresh; verify the model selector shows 智谱 GLM-4V-Flash(视觉).
Compatibility
DeepSeek Harness with skill support; Python 3.9+; vision model route GLM-4V-Flash (free) + DeepSeek-V4-Flash-Vision-Exp (same key as agent on v0.1.1) with failover chain to SiliconFlow Qwen3-VL → SenseNova → Google Gemini → any OpenAI-compatible endpoint. Paste-image reading requires a Harness build with core api-proxy image-admission support (bundled patches for rc.7/rc.8/v0.1.1-rc.1).
Details
- Repo: MJorgin/dsh-media-skills
- Category: Agent Capabilities
- Stars: 14
- Version: v0.1.1-rc.1 era (README badges: Harness rc.7 / rc.8 / v0.1.1 rc.1); MIT
- Last push: 2026-09-21
- First seen: 2026-08-14
Recent updates
The current README documents the two skills (vision-review, media-tools), the engine failover chain, the zero-extra-key v0.1.1 path vs rc.7/rc.8 patched path, key storage locations, the ModLens comparison/coexistence, and the 9-language docs.
FAQ
- Does paste-image reading need a Harness core patch?
- The auto-describe pipeline lives in Harness core (api-proxy image-admission logic); the bundle ships the model route + skills. The vision model works on any DSH build, but paste reading requires a build with that core support (patches bundled for rc.7/rc.8/v0.1.1-rc.1).
- Is media-tools really free?
- Yes — SiliconFlow Kolors is free and watermark-free; if a model is temporarily disabled the skill lists available models and you can switch.
- Where are API keys stored?
- Never in the repo: skill scripts read env vars → ~/.dsh/secrets/media-tools.env → ~/.codex/secrets/media-tools.env, and the vision route reads GLM_API_KEY from DSH's credential store.
Alternatives
william-jin-cmu/dsh-vision · oil-oil/dsh-vision · liustack/modlens