ruby1304/dsh-vision-subagent
Vision for text-only agents: a vision_agent tool that delegates image reading to a one-shot subagent on a configurable vision route (MiniMax/Kimi), plus a Codex-style paste bridge — images are analyzed on an isolated context and only text reaches the main session.
dsh-vision-subagent gives text-only DeepSeek Harness agents eyes by delegating image reading to a one-shot subagent running on a separately configured vision route (MiniMax, Kimi, or any OpenAI-compatible provider). Context isolation is the point: large screenshots and multi-image comparisons never occupy the main model's window, the child can call read_image on more workspace files before answering (multi-turn visual reasoning), and vision calls bill on the MiniMax/Kimi route while the main model only reasons. Pasting or dropping an image into the Web composer works Codex-style: the client uploads the image to the host endpoint, which stores it as a durable attachment and runs ONE vision-route analysis on an isolated context guided by your draft message — image bytes and intermediate context never enter the main session, only the final text answer returns.
Install
dsh plugin --profile web add dsh-vision-subagentnpm dsh-vision-subagent 0.3.0 verified 2026-09-04 (repository field → github.com/ruby1304/dsh-vision-subagent; README EN primary). Install: dsh plugin --profile web add dsh-vision-subagent, then add the vision-route row to ~/.dsh/profiles/web/cordis.patch.yml (e.g. provider: kimi-coding / model: k3, or minimax-cn / MiniMax-VL-01), restart dsh web, and ask the agent to look at an image. Image bytes and the vision model's intermediate context never enter the main session — only the final text answer comes back.
Compatibility
DeepSeek Harness Web (0.1.x per README); delegates to a one-shot subagent on a separately configured vision route (MiniMax / Kimi / any OpenAI-compatible provider); paste/drop image intake is native.
Details
- Repo: ruby1304/dsh-vision-subagent
- Category: Memory
- Stars: 0
- Version: npm dsh-vision-subagent 0.3.0
- Last push: 2026-09-10
- First seen: 2026-08-15
Recent updates
0.3.0 is the current npm latest (verified 2026-09-04).
FAQ
- Which vision routes can it use?
- A separately configured vision route — MiniMax, Kimi, or any OpenAI-compatible provider — set in cordis.patch.yml; the main session's model stays text-only.
- Do image bytes enter the main session?
- No — images are analyzed in an isolated subagent context and only the final text answer comes back; the main model never receives the raw image or the vision model's intermediate reasoning.
- Can it do multi-turn visual reasoning?
- Yes — the child subagent can call read_image on additional workspace files before answering, which is what makes comparisons and follow-up checks possible.
Alternatives
niuniuaba/dsh-subagent-vision · 1710782766/dsh-llm-vision