diff dictation/README.md @ 275:78699f810817

Add Qwen3-VL and WebRTC dictation services Add Bazel targets for the CUDA-backed Qwen3-VL server and a local WebRTC faster-whisper dictation service. Co-authored-by: Copilot <[email protected]> Copilot-Session: e3d8cb06-6c95-4ae0-9757-651d3796ab00
author MrJuneJune <me@mrjunejune.com>
date Mon, 17 Aug 2026 10:58:47 -0700
parents
children 49e9e591c9bb
line wrap: on
line diff
--- /dev/null	Thu Jan 01 00:00:00 1970 +0000
+++ b/dictation/README.md	Mon Aug 17 10:58:47 2026 -0700
@@ -0,0 +1,46 @@
+# WebRTC dictation
+
+This Bazel package runs a local WebRTC speech-to-text service using aiortc and
+faster-whisper. The default multilingual Whisper `small` model uses the RTX
+4070 Ti through native WSL CUDA with `int8_float16` compute.
+
+## Setup
+
+```bash
+bazel run //dictation:preflight
+bazel run //dictation:download_model
+bazel run //dictation:server
+```
+
+Open <http://127.0.0.1:8090>, grant microphone permission, and start
+dictation. Audio is carried by WebRTC and partial/final transcript events are
+returned on the `transcripts` RTC data channel.
+
+To transcribe an existing audio file:
+
+```bash
+bazel run //dictation:transcribe -- /path/to/audio.wav
+```
+
+To exercise the running HTTP signaling and WebRTC path with an audio file:
+
+```bash
+bazel run //dictation:webrtc_smoke -- /path/to/audio.wav
+```
+
+The model is stored under `~/.cache/zenbu/faster-whisper-small` by default.
+Override it with `DICTATION_MODEL_DIR`.
+
+Useful tuning:
+
+```bash
+DICTATION_COMPUTE_TYPE=int8_float16 \
+DICTATION_MAX_SESSIONS=1 \
+DICTATION_PARTIAL_INTERVAL_MS=1200 \
+DICTATION_SILENCE_MS=700 \
+  bazel run //dictation:server
+```
+
+The first version intentionally binds to loopback and does not configure STUN
+or TURN. Remote WebRTC access requires HTTPS, authentication, origin policy,
+and usually a TURN service.