WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866918092568788992 |
|---|---|
| author | Orhon, Atila Okan, Arda Durmus, Berkin Nagengast, Zach Pacheco, Eduardo |
| author_facet | Orhon, Atila Okan, Arda Durmus, Berkin Nagengast, Zach Pacheco, Eduardo |
| contents | Real-time Automatic Speech Recognition (ASR) is a fundamental building block for many commercial applications of ML, including live captioning, dictation, meeting transcriptions, and medical scribes. Accuracy and latency are the most important factors when companies select a system to deploy. We present WhisperKit, an optimized on-device inference system for real-time ASR that significantly outperforms leading cloud-based systems. We benchmark against server-side systems that deploy a diverse set of models, including a frontier model (OpenAI gpt-4o-transcribe), a proprietary model (Deepgram nova-3), and an open-source model (Fireworks large-v3-turbo).Our results show that WhisperKit matches the lowest latency at 0.46s while achieving the highest accuracy 2.2% WER. The optimizations behind the WhisperKit system are described in detail in this paper. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_10860 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | WhisperKit: On-device Real-time ASR with Billion-Scale Transformers Orhon, Atila Okan, Arda Durmus, Berkin Nagengast, Zach Pacheco, Eduardo Sound Computation and Language Audio and Speech Processing Real-time Automatic Speech Recognition (ASR) is a fundamental building block for many commercial applications of ML, including live captioning, dictation, meeting transcriptions, and medical scribes. Accuracy and latency are the most important factors when companies select a system to deploy. We present WhisperKit, an optimized on-device inference system for real-time ASR that significantly outperforms leading cloud-based systems. We benchmark against server-side systems that deploy a diverse set of models, including a frontier model (OpenAI gpt-4o-transcribe), a proprietary model (Deepgram nova-3), and an open-source model (Fireworks large-v3-turbo).Our results show that WhisperKit matches the lowest latency at 0.46s while achieving the highest accuracy 2.2% WER. The optimizations behind the WhisperKit system are described in detail in this paper. |
| title | WhisperKit: On-device Real-time ASR with Billion-Scale Transformers |
| topic | Sound Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2507.10860 |