Mercurial
view dictation/README.md @ 280:49e9e591c9bb
Add persistent dictation, prewarmed WebRTC speech input, Copilot SDK routing, animated conversation lifecycle controls, parking, and architecture coverage.
| author | MrJuneJune <me@mrjunejune.com> |
|---|---|
| date | Tue, 18 Aug 2026 19:14:53 -0700 |
| parents | 78699f810817 |
| children |
line wrap: on
line source
# WebRTC dictation This Bazel package runs a local WebRTC speech-to-text service using aiortc and faster-whisper. The default multilingual Whisper `small` model uses the RTX 4070 Ti through native WSL CUDA with `int8_float16` compute. ## Setup ```bash bazel run //dictation:preflight bazel run //dictation:download_model bazel run //dictation:server ``` Open <http://127.0.0.1:8090>, grant microphone permission, and start dictation. Audio is carried by WebRTC and partial/final transcript events are returned on the `transcripts` RTC data channel. To transcribe an existing audio file: ```bash bazel run //dictation:transcribe -- /path/to/audio.wav ``` To exercise the running HTTP signaling and WebRTC path with an audio file: ```bash bazel run //dictation:webrtc_smoke -- /path/to/audio.wav ``` The model is stored under `~/.cache/zenbu/faster-whisper-small` by default. Override it with `DICTATION_MODEL_DIR`. Useful tuning: ```bash DICTATION_COMPUTE_TYPE=int8_float16 \ DICTATION_MAX_SESSIONS=1 \ DICTATION_PARTIAL_INTERVAL_MS=500 \ DICTATION_SILENCE_MS=400 \ bazel run //dictation:server ``` The first version intentionally binds to loopback and does not configure STUN or TURN. Remote WebRTC access requires HTTPS, authentication, origin policy, and usually a TURN service. Infinite Canvas hosts this page as a hidden CEF/iframe transport. Pressing `M` shows partial and final transcript text in a centered Raylib caption; pressing `Enter` sends the accumulated text to Qwen session orchestration: ```bash bazel run //infinite_canvas:agent_dev ```