Whispy: Adapting STT Whisper Models to Real-Time Environments

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bevilacqua, Antonio, Saviano, Paolo, Amirante, Alessandro, Romano, Simon Pietro
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929336416731136
author Bevilacqua, Antonio
Saviano, Paolo
Amirante, Alessandro
Romano, Simon Pietro
author_facet Bevilacqua, Antonio
Saviano, Paolo
Amirante, Alessandro
Romano, Simon Pietro
contents Large general-purpose transformer models have recently become the mainstay in the realm of speech analysis. In particular, Whisper achieves state-of-the-art results in relevant tasks such as speech recognition, translation, language identification, and voice activity detection. However, Whisper models are not designed to be used in real-time conditions, and this limitation makes them unsuitable for a vast plethora of practical applications. In this paper, we introduce Whispy, a system intended to bring live capabilities to the Whisper pretrained models. As a result of a number of architectural optimisations, Whispy is able to consume live audio streams and generate high level, coherent voice transcriptions, while still maintaining a low computational cost. We evaluate the performance of our system on a large repository of publicly available speech datasets, investigating how the transcription mechanism introduced by Whispy impacts on the Whisper output. Experimental results show how Whispy excels in robustness, promptness, and accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03484
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Whispy: Adapting STT Whisper Models to Real-Time Environments
Bevilacqua, Antonio
Saviano, Paolo
Amirante, Alessandro
Romano, Simon Pietro
Sound
Machine Learning
Audio and Speech Processing
Large general-purpose transformer models have recently become the mainstay in the realm of speech analysis. In particular, Whisper achieves state-of-the-art results in relevant tasks such as speech recognition, translation, language identification, and voice activity detection. However, Whisper models are not designed to be used in real-time conditions, and this limitation makes them unsuitable for a vast plethora of practical applications. In this paper, we introduce Whispy, a system intended to bring live capabilities to the Whisper pretrained models. As a result of a number of architectural optimisations, Whispy is able to consume live audio streams and generate high level, coherent voice transcriptions, while still maintaining a low computational cost. We evaluate the performance of our system on a large repository of publicly available speech datasets, investigating how the transcription mechanism introduced by Whispy impacts on the Whisper output. Experimental results show how Whispy excels in robustness, promptness, and accuracy.
title Whispy: Adapting STT Whisper Models to Real-Time Environments
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2405.03484