perrylink/dsh-talk
Voice-first conversation closed loop: composer microphone button + browser/local speech transcription (Web Speech, FunASR, whisper.cpp), speak tool to read replies (browser, edge-tts, piper), event broadcast with mute switch, speech interruption
dsh-talk closes the voice loop for DeepSeek Harness in both directions. The agent gets a speak tool with pluggable TTS engines (browser voice, edge-tts neural voices, or local piper); a composer microphone button captures speech and lands the transcription in the input box or submits directly, with STT engines ranging from the browser's Web Speech (interim results included) to a FunASR HTTP server or local whisper.cpp. Speak-to-interrupt stops whatever is playing the moment you start talking (client → host over the talk Remote namespace). Event announcements cover turn completion, pending approvals (waterfall-safe, never blocks the gate) and errors, with a mute switch and configurable phrases. A Settings tab selects engines/languages and announcement switches, saved as append-only profile-patch operations with backups.
Install
dsh plugin --profile web add dsh-talknpm dsh-talk 0.3.3 verified 2026-09-04 (repository field → github.com/PerryLink/dsh-talk; README EN primary with zh/es/pt/hi editions). Install: dsh plugin --profile web add dsh-talk (npm) or dsh plugin --profile web add "github:PerryLink/dsh-talk#main" (git). After install restart and verify: dsh --profile web --dump-config | grep -A2 'id: talk'.
Compatibility
DeepSeek Harness 0.1.2-alpha.5 (per README); Node ^22.19.0 || >=24.0.0; browser Web Speech + MediaRecorder (Chrome/Edge best); host transcription/TTS engines for the rest (edge-tts, piper, FunASR HTTP, local whisper.cpp).
Details
- Repo: perrylink/dsh-talk
- Category: Other
- Stars: 0
- Version: npm dsh-talk 0.3.3
- Last push: 2026-08-21
- First seen: 2026-08-16
Recent updates
0.3.3 is the current npm latest (verified 2026-09-04).
FAQ
- Which TTS engines are supported?
- Three: the browser voice, edge-tts (network neural voices) and piper (local); audio plays in the browser, and where the host supports it the session log records the sanitized utterance.
- Can I interrupt a spoken reply?
- Yes — speak-to-interrupt: starting to talk stops whatever is playing, via a client→host call over the talk Remote namespace.
- Does voice input need an API key?
- No — the browser's Web Speech engine works out of the box; FunASR or whisper.cpp are optional local upgrades, and edge-tts is the network neural TTS option.
Alternatives
stardustlc666/dsh-voice · Alan2Z/dsh-speak