sunshine-lang/dsh-pdf
PDF toolbox for DeepSeek Harness: extract text, metadata, and page ranges via pdfjs-dist (local, no API key)
A PDF toolbox for DeepSeek Harness: the agent can extract text content from local PDF files, read metadata (title, author, page count, creation date), and access specific page ranges — all processed locally without sending files to any external service. Three agent tools are registered: pdf_text (full text or page-range extraction), pdf_metadata (document properties), and pdf_pages (paginated text access for large documents).
Install
dsh plugin --profile web add dsh-pdfdsh plugin --profile web add dsh-pdf (npm dsh-pdf 0.1.0, repo sunshine-lang/dsh-pdf, verified npm 2026-08-30). PDF processing runs locally — no file is uploaded to any external service.
Compatibility
DSH web profile; local PDF parsing (no upload); pure Node.js implementation; MIT.
Details
- Repo: sunshine-lang/dsh-pdf
- Category: Files & Data
- Stars: 7
- Version: npm dsh-pdf 0.1.0
- Last push: 2026-09-21
- First seen: 2026-08-14
Recent updates
pdf_text (full or page-range extraction), pdf_metadata, pdf_pages tools; local processing only (no upload); metadata: title/author/page count/date.
FAQ
- Is the PDF parsed on the server or locally?
- Locally — the plugin uses a pure Node.js PDF parser running inside the DSH process; no file is sent to any external API or cloud service.
- What types of PDFs are supported?
- Text-layer PDFs (natively produced from Word, LaTeX, etc.) work reliably; scanned image-only PDFs require OCR which this plugin does not perform — for those, use a vision plugin or a dedicated OCR step first.
- Can the agent read very large PDFs?
- Yes — pdf_pages supports paginated access with a configurable page-range, so you can instruct the agent to read sections of a large document incrementally.
Alternatives
taxueseek/dsh-files · ben7am1n/dsh-lens-lite · zhangzujian/dsh-subprocess-inherit-environment