Demo: WhisperFlow: speech foundation models in real time

0Citations
Citations of this article
2Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Speech foundation models, such as OpenAI's Whisper, become the state of the art in speech understanding due to their strong accuracy and generalizability. Yet, their applications are mostly limited to processing pre-recorded speech, whereas processing of streaming speech, in particular doing it efficiently, remains rudimentary.We present a novel framework, WhisperFlow, which embodies both model and system optimizations. (1) Hush word as a short, learnable audio segment; appended to a voice input, a hush word gracefully stops the speech model from processing more input without hallucination; (2) Beam pruning, which aligns streaming audio buffers over time and reuses results from earlier decoding rounds, therefore significantly accelerating decoding; and (3) CPU/GPU pipelining, which not only maps to the encoding/decoding stages dynamically, but also tunes to an optimal resource ratio, respecting the encoding/decoding speed that varies across voice inputs, models, and hardware.We demonstrate WhisperFlow on a Macbook pro with M4 pro SoC with 14 CPU cores and 20 GPU cores on real-world conversation transcription tasks. The WhisperFlow delivers high fidelity transcripts with the Whisper medium model, also maintains the per-word latency within 1 second.

Cite

CITATION STYLE

APA

Wang, R., & Lin, F. X. (2025). Demo: WhisperFlow: speech foundation models in real time. In MobiSys 2025 - Proceedings of the 23rd ACM international Conference on Mobile Systems, Applications, and Services (pp. 621–622). Association for Computing Machinery, Inc. https://doi.org/10.1145/3711875.3734370

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free