PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915074660106240 |
|---|---|
| author | Nassereldine, Amir Liu, Dancheng Xu, Chenhui Qin, Ruiyang Shi, Yiyu Xiong, Jinjun |
| author_facet | Nassereldine, Amir Liu, Dancheng Xu, Chenhui Qin, Ruiyang Shi, Yiyu Xiong, Jinjun |
| contents | Edge-based automatic speech recognition (ASR) technologies are increasingly prevalent in the development of intelligent and personalized assistants. However, resource-constrained ASR models face significant challenges in adaptivity, incrementality, and inclusivity when faced with a diverse population. To tackle those challenges, we propose PI-Whisper, a novel ASR system that adaptively enhances recognition capabilities by identifying speakers' characteristics in real-time. In this work, we show how the design of PI-Whisper allows for incremental adaptation of new characteristics without the need for repetitive retraining, enhances recognition capabilities, and improves equity and fairness across diverse speaker groups. PI-Whisper demonstrates these advantages by achieving state-of-the-art accuracy, reducing the word error rate (WER) by up to 13.7% relative to baselines while scaling linearly to computing resources. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_15668 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices Nassereldine, Amir Liu, Dancheng Xu, Chenhui Qin, Ruiyang Shi, Yiyu Xiong, Jinjun Computation and Language Sound Audio and Speech Processing Edge-based automatic speech recognition (ASR) technologies are increasingly prevalent in the development of intelligent and personalized assistants. However, resource-constrained ASR models face significant challenges in adaptivity, incrementality, and inclusivity when faced with a diverse population. To tackle those challenges, we propose PI-Whisper, a novel ASR system that adaptively enhances recognition capabilities by identifying speakers' characteristics in real-time. In this work, we show how the design of PI-Whisper allows for incremental adaptation of new characteristics without the need for repetitive retraining, enhances recognition capabilities, and improves equity and fairness across diverse speaker groups. PI-Whisper demonstrates these advantages by achieving state-of-the-art accuracy, reducing the word error rate (WER) by up to 13.7% relative to baselines while scaling linearly to computing resources. |
| title | PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2406.15668 |