Mercurial
view dictation/README.md @ 275:78699f810817
Add Qwen3-VL and WebRTC dictation services
Add Bazel targets for the CUDA-backed Qwen3-VL server and a local WebRTC faster-whisper dictation service.
Co-authored-by: Copilot <[email protected]>
Copilot-Session: e3d8cb06-6c95-4ae0-9757-651d3796ab00
| author | MrJuneJune <me@mrjunejune.com> |
|---|---|
| date | Mon, 17 Aug 2026 10:58:47 -0700 |
| parents | |
| children | 49e9e591c9bb |
line wrap: on
line source
# WebRTC dictation This Bazel package runs a local WebRTC speech-to-text service using aiortc and faster-whisper. The default multilingual Whisper `small` model uses the RTX 4070 Ti through native WSL CUDA with `int8_float16` compute. ## Setup ```bash bazel run //dictation:preflight bazel run //dictation:download_model bazel run //dictation:server ``` Open <http://127.0.0.1:8090>, grant microphone permission, and start dictation. Audio is carried by WebRTC and partial/final transcript events are returned on the `transcripts` RTC data channel. To transcribe an existing audio file: ```bash bazel run //dictation:transcribe -- /path/to/audio.wav ``` To exercise the running HTTP signaling and WebRTC path with an audio file: ```bash bazel run //dictation:webrtc_smoke -- /path/to/audio.wav ``` The model is stored under `~/.cache/zenbu/faster-whisper-small` by default. Override it with `DICTATION_MODEL_DIR`. Useful tuning: ```bash DICTATION_COMPUTE_TYPE=int8_float16 \ DICTATION_MAX_SESSIONS=1 \ DICTATION_PARTIAL_INTERVAL_MS=1200 \ DICTATION_SILENCE_MS=700 \ bazel run //dictation:server ``` The first version intentionally binds to loopback and does not configure STUN or TURN. Remote WebRTC access requires HTTPS, authentication, origin policy, and usually a TURN service.