Elohia/pi-mm-vision
A 'synesthesia encoder' that lets text-only LLMs see images: in one API call, a vision model translates an image into a compact, coordinate-structured spatial description (canvas, elements, (x%, y%) positions, shapes, values, relations) that a text-only model can rebuild and reason about at pixel level. Four modes (brief/full/coords/auto), an optional dot-matrix mode for curve detail, a TTL cache (default 600s), and a two-way mcp_render protocol that turns the structured description back into SVG so a text-only LLM can draw. In DSH it registers as the mm_vision model tool.
Install
dsh plugin --profile web add dsh-plugin-mm-visiondsh plugin --profile web add dsh-plugin-mm-vision (npm dsh-plugin-mm-vision 0.1.1, verified npm 2026-08-31; repo Elohia/pi-mm-vision). GitHub direct install also works: dsh plugin add github:Elohia/dsh-plugin-mm-vision. Multi-host project — the same encoder also ships as MCP, CLI, and agent extensions.
Compatibility
DSH web profile; works with any OpenAI-compatible vision model (qwen-vl, gpt-4o, glm-4v, kimi-vl, MiniMax-VL, …); model/baseUrl/API key fully configurable; MIT.
Details
- Repo: Elohia/pi-mm-vision
- Category: Other
- Stars: 2
- Version: npm dsh-plugin-mm-vision 0.1.1
- Last push: 2026-08-14
- First seen: 2026-08-03
Recent updates
Coordinate-structured image encoding (canvas/elements/percent coords); brief/full/coords/auto modes; optional dotMatrix; 600s TTL cache; mcp_render → SVG drawing; DSH mm_vision tool.
FAQ
- How is this different from ASCII art or natural-language description?
- ASCII is slow and token-heavy with no colour or semantics; prose loses geometry. The encoder uses one API call to produce compact coordinate text so position reasoning stays pixel-accurate.
- Which vision models work?
- Any OpenAI-compatible vision endpoint — qwen-vl, gpt-4o, glm-4v, kimi-vl, MiniMax-VL, etc. — with model, baseUrl, and API key all configurable.
- Can the text-only model draw images?
- Yes — the two-way mcp_render protocol converts the structured encoding into SVG (rectangles, circles, polygons, arrows, text, terrain), so the LLM gains a drawing capability.
Alternatives
RRRosmontis/dsh-qwen-mm · Yuuz12/dsh-vision-helper · 54xkeee/dsh-youreyes