maxwell-feng/dsh-tesseract-ocr

Local OCR for attached images via Tesseract: only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.

dsh-tesseract-ocr lets text-only DeepSeek Harness models accept attached images: every image is recognized locally with Tesseract OCR and only the recognized text is sent to the model API — image bytes never leave the machine by default. Vision-model passthrough is opt-in (passthrough: true) for models that genuinely accept original image bytes. The plugin hooks two public llm seams: a capability shim (resolveModelInfo/listModels) so text models admit image attachments, and an agent/pre-step rewrite that replaces every image content block with recognized text before the request leaves the machine. Fail-closed: if the plugin is not loaded, models stay text-only and image attachments are refused; missing attachments are replaced with a refusal text block, never left as raw image.

Vision & Multimodal ★ 0 updated 2026-09-13

Source repository unavailable. The repository maxwell-feng/dsh-tesseract-ocr no longer exists on GitHub (HTTP 404 when checked on 2026-10-02). The plugin is still published on npm as dsh-tesseract-ocr, which is the install command shown below.

Repository unavailable ↗

Install

dsh plugin --profile web add dsh-tesseract-ocr

npm dsh-tesseract-ocr 0.3.7 verified 2026-09-04 (repository field → github.com/maxwell-feng/dsh-tesseract-ocr; README EN primary). Install: dsh plugin --profile web add dsh-tesseract-ocr (replace web with your profile). Requires the tesseract CLI installed on the host (Linux/macOS/Windows); the npm package ships prebuilt with Sigstore provenance — no allowBuilds approval needed. Do NOT enable together with dsh-windows-ocr: both would OCR the same image — pick one per machine.

Compatibility

DeepSeek Harness (verified against 0.1.2-rc.1 master per README); works anywhere the tesseract CLI is installed (Linux, macOS, Windows); any provider/model in dsh.

Details

Recent updates

0.3.7 is the current npm latest (verified 2026-09-04).

FAQ

Are image bytes ever sent to the provider?
By default no — OCR runs locally with the Tesseract CLI and only recognized text is sent; set passthrough: true only if you intentionally want a genuine vision model to receive original image bytes.
Do I need to change model config?
No — no input: [text, image] hacks in settings.yaml; the capability shim answers 'yes' at the host's image-admission checks, so text models accept attachments while the plugin is loaded.
What happens if the plugin is not loaded?
Fail-closed: models stay text-only and image attachments are refused — nothing can silently leak; a missing attachment is replaced with a refusal text block.

Alternatives

Aidenwu0209/dsh-Unlimited-OCR-Skill · good-boy4069/dsh-vision-guard

More plugins in Vision & Multimodal

Browse more in Vision & Multimodal

Guides for Vision & Multimodal plugins