Real-time Whisper on-device: chunk size, VAD, and quantization, with numbers
Benchmarking whisper.cpp for live meeting transcription on Apple Silicon: the ~1.6s latency floor, why VAD is a 28% win, why q5_0 is free, and why CoreML wasn't worth it.
Read