Zhangbo-cn/dsh-voice-input-plugin

Composer mic for the Web UI: tap-to-monitor live transcription and hold-to-talk, with host Edge TTS reply reading that streams while the model generates, echo-pause during reading, and tap-to-stop.

Adds a minimal linear mic button to the DSH composer tool row with two voice modes: tap-to-monitor (continuous live streaming dictation, send-anytime, auto-send on silence is optional) and hold-to-talk voice chat (release to send, assistant reply read aloud). Each recognition segment auto-restarts so monitoring never drops, the icon pulses DeepSeek-blue while listening, and speech appends to the existing draft — base text preserved, draft cleared cleanly after send.

UI Enhancements ★ 6 updated 2026-08-25
View on GitHub ↗

Install

dsh plugin --profile web add @zhangbo-cn/dsh-client-ui-voice-input

dsh plugin --profile web add @zhangbo-cn/dsh-client-ui-voice-input (npm @zhangbo-cn/dsh-client-ui-voice-input, repo Zhangbo-cn/dsh-voice-input-plugin, verified npm 2026-08-31). The README requires 0.1.1+ — 0.1.0 registered the browser bundle under the wrong ModuleLoader id; if your registry resolves 0.1.0, install from GitHub (dsh plugin --profile web add github:Zhangbo-cn/dsh-voice-input-plugin) instead.

Compatibility

DSH web profile; recognition runs in-browser via the Web Speech API (zero API key); reply reading uses host Edge TTS (/api/tts) with browser speechSynthesis fallback; TypeScript/React; MIT.

Details

Recent updates

Tap-to-monitor + hold-to-talk voice modes; optional auto-send on silence; host Edge TTS with speechSynthesis fallback; continuous recognition across silences; configurable recognition language (default zh-CN) and interim results.

FAQ

Does voice recognition need an API key?
No — recognition runs entirely in the browser via the Web Speech API. Reply reading uses the host's Edge TTS endpoint with a browser speechSynthesis fallback.
What is tap-to-monitor mode?
Click the mic and speak — text streams into the draft live and the mic keeps listening even in silence; tap again to stop. An optional autoSendOnSilenceMs setting makes it hands-free.
What does hold-to-talk do?
Press and hold to record a voice-chat message; release to send it. The assistant's reply is then read aloud via host neural TTS.

Alternatives

3274375092/dsh-voice · forrestahha/dsh-voice-input · Jesse-njx/dsh-voice

More plugins in UI Enhancements

Browse more in UI Enhancements

Guides for UI Enhancements plugins