sunshine-lang/dsh-pdf

PDF toolbox for DeepSeek Harness: extract text, metadata, and page ranges via pdfjs-dist (local, no API key)

A PDF toolbox for DeepSeek Harness: the agent can extract text content from local PDF files, read metadata (title, author, page count, creation date), and access specific page ranges — all processed locally without sending files to any external service. Three agent tools are registered: pdf_text (full text or page-range extraction), pdf_metadata (document properties), and pdf_pages (paginated text access for large documents).

Files & Data ★ 7 updated 2026-09-21 ✅ runtime-tested
View on GitHub ↗

Install

dsh plugin --profile web add dsh-pdf

dsh plugin --profile web add dsh-pdf (npm dsh-pdf 0.1.0, repo sunshine-lang/dsh-pdf, verified npm 2026-08-30). PDF processing runs locally — no file is uploaded to any external service.

Compatibility

DSH web profile; local PDF parsing (no upload); pure Node.js implementation; MIT.

Details

Recent updates

pdf_text (full or page-range extraction), pdf_metadata, pdf_pages tools; local processing only (no upload); metadata: title/author/page count/date.

FAQ

Is the PDF parsed on the server or locally?
Locally — the plugin uses a pure Node.js PDF parser running inside the DSH process; no file is sent to any external API or cloud service.
What types of PDFs are supported?
Text-layer PDFs (natively produced from Word, LaTeX, etc.) work reliably; scanned image-only PDFs require OCR which this plugin does not perform — for those, use a vision plugin or a dedicated OCR step first.
Can the agent read very large PDFs?
Yes — pdf_pages supports paginated access with a configurable page-range, so you can instruct the agent to read sections of a large document incrementally.

Alternatives

taxueseek/dsh-files · ben7am1n/dsh-lens-lite · zhangzujian/dsh-subprocess-inherit-environment

More plugins in Files & Data

Browse more in Files & Data

Guides for Files & Data plugins