maxwell-feng/dsh-windows-ocr

Local OCR for attached images via the built-in Windows engine (Windows.Media.Ocr): only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.

dsh-windows-ocr is the Windows-native sibling of dsh-tesseract-ocr: it lets text-only DeepSeek Harness models accept attached images by recognizing every image locally with the built-in Windows OCR engine (Windows.Media.Ocr) and sending only the recognized text to the model API. Privacy default is local-only OCR — image bytes are not sent to the provider unless you opt into passthrough: true for genuine vision models. It hooks the same two public llm seams (capability shim + agent/pre-step rewrite) and is fail-closed: without the plugin loaded, models stay text-only and image attachments are refused.

Vision & Multimodal ★ 0 updated 2026-09-13

Source repository unavailable. The repository maxwell-feng/dsh-windows-ocr no longer exists on GitHub (HTTP 404 when checked on 2026-10-02). The plugin is still published on npm as dsh-windows-ocr, which is the install command shown below.

Repository unavailable ↗

Install

dsh plugin --profile web add dsh-windows-ocr

npm dsh-windows-ocr 0.3.7 verified 2026-09-04 (repository field → github.com/maxwell-feng/dsh-windows-ocr; README EN primary). Install: dsh plugin --profile web add dsh-windows-ocr. Windows-only: uses the built-in Windows.Media.Ocr engine, so no separate OCR binary is needed; the npm package ships prebuilt with Sigstore provenance. Do NOT enable together with dsh-tesseract-ocr — both would OCR the same image; pick one per machine.

Compatibility

Windows + DeepSeek Harness; uses the built-in Windows OCR engine (Windows.Media.Ocr); any provider/model in dsh; no extra OCR install.

Details

Recent updates

0.3.7 is the current npm latest (verified 2026-09-04).

FAQ

Is this the same as dsh-tesseract-ocr?
Same design and fail-closed behavior, but it uses Windows' built-in OCR engine (Windows.Media.Ocr) instead of the Tesseract CLI — no separate binary to install. Do not run both plugins at once.
Which Windows versions work?
Any Windows where the Windows.Media.Ocr API is available (Windows 10/11); the README tests primarily on Windows with the built-in engine.
Are images uploaded anywhere?
No — recognition happens locally on Windows and only the recognized text is sent to the model, unless you explicitly set passthrough: true for a genuine vision model.

Alternatives

Aidenwu0209/dsh-PaddleOCR-Skills · maxwell-feng/dsh-tesseract-ocr

More plugins in Vision & Multimodal

Browse more in Vision & Multimodal

Guides for Vision & Multimodal plugins