WhisperKit: On-device Real-time ASR with Billion-Scale Transformers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Orhon, Atila, Okan, Arda, Durmus, Berkin, Nagengast, Zach, Pacheco, Eduardo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918092568788992
author Orhon, Atila
Okan, Arda
Durmus, Berkin
Nagengast, Zach
Pacheco, Eduardo
author_facet Orhon, Atila
Okan, Arda
Durmus, Berkin
Nagengast, Zach
Pacheco, Eduardo
contents Real-time Automatic Speech Recognition (ASR) is a fundamental building block for many commercial applications of ML, including live captioning, dictation, meeting transcriptions, and medical scribes. Accuracy and latency are the most important factors when companies select a system to deploy. We present WhisperKit, an optimized on-device inference system for real-time ASR that significantly outperforms leading cloud-based systems. We benchmark against server-side systems that deploy a diverse set of models, including a frontier model (OpenAI gpt-4o-transcribe), a proprietary model (Deepgram nova-3), and an open-source model (Fireworks large-v3-turbo).Our results show that WhisperKit matches the lowest latency at 0.46s while achieving the highest accuracy 2.2% WER. The optimizations behind the WhisperKit system are described in detail in this paper.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10860
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
Orhon, Atila
Okan, Arda
Durmus, Berkin
Nagengast, Zach
Pacheco, Eduardo
Sound
Computation and Language
Audio and Speech Processing
Real-time Automatic Speech Recognition (ASR) is a fundamental building block for many commercial applications of ML, including live captioning, dictation, meeting transcriptions, and medical scribes. Accuracy and latency are the most important factors when companies select a system to deploy. We present WhisperKit, an optimized on-device inference system for real-time ASR that significantly outperforms leading cloud-based systems. We benchmark against server-side systems that deploy a diverse set of models, including a frontier model (OpenAI gpt-4o-transcribe), a proprietary model (Deepgram nova-3), and an open-source model (Fireworks large-v3-turbo).Our results show that WhisperKit matches the lowest latency at 0.46s while achieving the highest accuracy 2.2% WER. The optimizations behind the WhisperKit system are described in detail in this paper.
title WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2507.10860