Mercurial
view dictation/README.md @ 279:b3b547563ec7
Add Google connector service and agent wiki
Implement the C/Seobeo Google Drive and Gmail connector with encrypted OAuth storage, Zenbu authentication, browser testing, AI tool discovery, chunked HTTP decoding, and Bazel coverage. Consolidate repository guidance into progressive wiki documentation and enforce arena-first allocation for new first-party C code.
Co-authored-by: Copilot <[email protected]>
Copilot-Session: 84c338fd-0939-4bb3-b7f3-1062eb213e5d
| author | MrJuneJune <me@mrjunejune.com> |
|---|---|
| date | Mon, 17 Aug 2026 22:22:36 -0700 |
| parents | 78699f810817 |
| children | 49e9e591c9bb |
line wrap: on
line source
# WebRTC dictation This Bazel package runs a local WebRTC speech-to-text service using aiortc and faster-whisper. The default multilingual Whisper `small` model uses the RTX 4070 Ti through native WSL CUDA with `int8_float16` compute. ## Setup ```bash bazel run //dictation:preflight bazel run //dictation:download_model bazel run //dictation:server ``` Open <http://127.0.0.1:8090>, grant microphone permission, and start dictation. Audio is carried by WebRTC and partial/final transcript events are returned on the `transcripts` RTC data channel. To transcribe an existing audio file: ```bash bazel run //dictation:transcribe -- /path/to/audio.wav ``` To exercise the running HTTP signaling and WebRTC path with an audio file: ```bash bazel run //dictation:webrtc_smoke -- /path/to/audio.wav ``` The model is stored under `~/.cache/zenbu/faster-whisper-small` by default. Override it with `DICTATION_MODEL_DIR`. Useful tuning: ```bash DICTATION_COMPUTE_TYPE=int8_float16 \ DICTATION_MAX_SESSIONS=1 \ DICTATION_PARTIAL_INTERVAL_MS=1200 \ DICTATION_SILENCE_MS=700 \ bazel run //dictation:server ``` The first version intentionally binds to loopback and does not configure STUN or TURN. Remote WebRTC access requires HTTPS, authentication, origin policy, and usually a TURN service.