PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nassereldine, Amir, Liu, Dancheng, Xu, Chenhui, Qin, Ruiyang, Shi, Yiyu, Xiong, Jinjun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915074660106240
author Nassereldine, Amir
Liu, Dancheng
Xu, Chenhui
Qin, Ruiyang
Shi, Yiyu
Xiong, Jinjun
author_facet Nassereldine, Amir
Liu, Dancheng
Xu, Chenhui
Qin, Ruiyang
Shi, Yiyu
Xiong, Jinjun
contents Edge-based automatic speech recognition (ASR) technologies are increasingly prevalent in the development of intelligent and personalized assistants. However, resource-constrained ASR models face significant challenges in adaptivity, incrementality, and inclusivity when faced with a diverse population. To tackle those challenges, we propose PI-Whisper, a novel ASR system that adaptively enhances recognition capabilities by identifying speakers' characteristics in real-time. In this work, we show how the design of PI-Whisper allows for incremental adaptation of new characteristics without the need for repetitive retraining, enhances recognition capabilities, and improves equity and fairness across diverse speaker groups. PI-Whisper demonstrates these advantages by achieving state-of-the-art accuracy, reducing the word error rate (WER) by up to 13.7% relative to baselines while scaling linearly to computing resources.
format Preprint
id arxiv_https___arxiv_org_abs_2406_15668
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
Nassereldine, Amir
Liu, Dancheng
Xu, Chenhui
Qin, Ruiyang
Shi, Yiyu
Xiong, Jinjun
Computation and Language
Sound
Audio and Speech Processing
Edge-based automatic speech recognition (ASR) technologies are increasingly prevalent in the development of intelligent and personalized assistants. However, resource-constrained ASR models face significant challenges in adaptivity, incrementality, and inclusivity when faced with a diverse population. To tackle those challenges, we propose PI-Whisper, a novel ASR system that adaptively enhances recognition capabilities by identifying speakers' characteristics in real-time. In this work, we show how the design of PI-Whisper allows for incremental adaptation of new characteristics without the need for repetitive retraining, enhances recognition capabilities, and improves equity and fairness across diverse speaker groups. PI-Whisper demonstrates these advantages by achieving state-of-the-art accuracy, reducing the word error rate (WER) by up to 13.7% relative to baselines while scaling linearly to computing resources.
title PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.15668