welsione/dsh-mmx-bridge

MiniMax multimodal bridge: one mmx_bridge tool covers image understanding/generation, video, TTS, music, cover, web search and quota; optional web_search/read_image takeover; inline players/image previews right in the Web GUI (npm: dsh-mmx-bridge).

dsh-mmx-bridge plugs MiniMax's full multimodal stack into DeepSeek Harness through a single mmx_bridge tool — 8 capabilities: image recognition (VLM describe/analyze, with embedded JSON cache in PNG tEXt / JPEG COM blocks so the same image+question is never re-sent), image generation (art), video creation, text-to-speech, song generation, web search, and more. Since v1.0.5, dropping or pasting an image into the chat input works even with text-only models: the plugin saves to a temp dir and replaces the image with 'image URL + local path' text; the agent then auto-calls read_image / mmx_bridge(describe). The recognition cache (v1.0.7+) writes results back into the image file as imgjson blocks; follow-up questions are cached per prompt layer and invalidate automatically when the image is re-encoded.

Tools & Capabilities ★ 0 updated 2026-09-18
View on GitHub ↗

Install

dsh plugin --profile web add dsh-mmx-bridge

The npm package dsh-mmx-bridge 1.0.10 was re-verified live on 2026-10-01 with its repository field back-linking to welsione/dsh-mmx-bridge. After installing, add your MiniMax API key under Settings → MiniMax.

Compatibility

DeepSeek Harness 0.1.0-rc.7+. Requires a MiniMax API key (set under Settings → MiniMax). Image recognition results are cached per image as embedded JSON (PNG tEXt / JPEG COM blocks) — same image + same question served from cache, zero VLM calls on second read.

Details

Recent updates

The README documents the 8 capabilities, the drop-and-paste flow for text-only models (v1.0.5+), the embedded image recognition cache (v1.0.7+), and the minimax-cli dependency.

FAQ

How do I install dsh-mmx-bridge?
dsh plugin --profile web add dsh-mmx-bridge (npm 1.0.10 re-verified 2026-10-01), then set your MiniMax API key under Settings → MiniMax.
Can I drop images even if my model is text-only?
Yes — since v1.0.5, the plugin saves dropped/pasted images to a temp dir and feeds the agent a URL+path; the agent then calls read_image/mmx_bridge to see it.
Does it re-process the same image every time?
No — since v1.0.7, recognition results are embedded in the image file as imgjson blocks; same image + same question is served from cache, zero VLM calls on later reads.

Alternatives

jiavenzhong/dsh-tool-markitdown · shixiliya1/dsh-rich-file-reader · whitefirer/dsh-browser-fs

More plugins in Tools & Capabilities

Browse more in Tools & Capabilities

Guides for Tools & Capabilities plugins