RRRosmontis/dsh-qwen-mm
Qwen-MM-Plugins integration bundle for DeepSeek Harness (dsh) — multimodal MCP tools (vision, OCR, ASR, search
Makes DeepSeek Harness multimodal-native by integrating Qwen-MM-Plugins: core MCP tools (read_image, visualize, media_info, read_video…) read images/videos/documents/code/data with no key; api MCP tools (vision_chat, ocr, omni_*) add cloud vision/OCR/ASR/object localization via DashScope Qwen VL/Omni; search MCP tools (web_search, web_extractor, image_search) add web + reverse-image search; video-memory / video-edit / blender / freecad / edu-agent extend to long-video QA, editing, 3D and math. A three-part image-attachment bridge lets you drag an image onto a pure-text model: the dsh-attachment consumer registry allows image upload on text-only routes, the attachments bridge plugin exports each image block to <dshHome>/qwen-mm/attachments/<sha256>.<ext> and rewrites the message to a path reference, and usage guidance tells the model to read it via mcpqwen-mm-plugins-apivision_chat or mcpqwen-mm-plugins-coreread_image.
Install
pnpm dsh --profile web plugin add github:RRRosmontis/dsh-qwen-mmDrag-image support needs a small DSH core change (image-consumer registry) not yet in official releases, so the README recommends installing from the maintained fork: git clone https://github.com/RRRosmontis/deepseek-harness.git && cd deepseek-harness && pnpm install && pnpm run build, then pnpm dsh --profile web plugin add github:RRRosmontis/dsh-qwen-mm (or ./packages/bundle/qwen-mm) and pnpm dsh --profile web. Prereqs: Node.js, pnpm, and uv (for MCP python bridges). Bilingual README with full English section.
Compatibility
DSH (fork with image-consumer registry for full drag-image support); pure-text DeepSeek models; needs DashScope key for cloud vision/OCR/ASR, Serper/Tavily/Exa for search tools. MCP-based.
Details
- Repo: RRRosmontis/dsh-qwen-mm
- Category: Coding & Development
- Stars: 2
- Version: GitHub source (bundle in fork; no standalone npm)
- Last push: 2026-08-13
- First seen: 2026-08-13
Recent updates
Qwen-MM-Plugins integration; keyless core vision tools; DashScope cloud vision/OCR/ASR; web + reverse-image search; drag-image bridge for pure-text models.
FAQ
- Which tools need a key?
- Core read/visualize/media tools are keyless; cloud vision/OCR/ASR needs a DashScope key; web search needs Serper/Tavily/Exa.
- Can pure-text DeepSeek models see images?
- Yes — the attachment bridge lets you drag an image in; the message is rewritten to a path reference and the model reads it via the Qwen VL MCP tools.
- Do I need the fork?
- For full drag-image support yes (it contains the image-consumer registry patch); core MCP tools work without it.
Alternatives
Scorp1o117/dsh-tool-vision · yuz12-dsh-vision-helper · libinyam/dsh-vision-provider