Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nagpal, Chirag, Venugopalan, Subhashini, Tobin, Jimmy, Ladewig, Marilyn, Heller, Katherine, Tomanek, Katrin
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915086680981504
author Nagpal, Chirag
Venugopalan, Subhashini
Tobin, Jimmy
Ladewig, Marilyn
Heller, Katherine
Tomanek, Katrin
author_facet Nagpal, Chirag
Venugopalan, Subhashini
Tobin, Jimmy
Ladewig, Marilyn
Heller, Katherine
Tomanek, Katrin
contents We introduce a large language model (LLM) capable of processing speech inputs and show that tuning it further with reinforcement learning on human preference (RLHF) enables it to adapt better to disordered speech than traditional fine-tuning. Our method replaces low-frequency text tokens in an LLM's vocabulary with audio tokens and enables the model to recognize speech by fine-tuning it on speech with transcripts. We then use RL with rewards based on syntactic and semantic accuracy measures generalizing the LLM further to recognize disordered speech. While the resulting LLM does not outperform existing systems for speech recognition, we find that tuning with reinforcement learning using custom rewards leads to substantially better performance than supervised fine-tuning of the language model, specifically when adapting to speech in a different setting. This presents a compelling alternative tuning strategy for speech recognition using large language models.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00039
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
Nagpal, Chirag
Venugopalan, Subhashini
Tobin, Jimmy
Ladewig, Marilyn
Heller, Katherine
Tomanek, Katrin
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
We introduce a large language model (LLM) capable of processing speech inputs and show that tuning it further with reinforcement learning on human preference (RLHF) enables it to adapt better to disordered speech than traditional fine-tuning. Our method replaces low-frequency text tokens in an LLM's vocabulary with audio tokens and enables the model to recognize speech by fine-tuning it on speech with transcripts. We then use RL with rewards based on syntactic and semantic accuracy measures generalizing the LLM further to recognize disordered speech. While the resulting LLM does not outperform existing systems for speech recognition, we find that tuning with reinforcement learning using custom rewards leads to substantially better performance than supervised fine-tuning of the language model, specifically when adapting to speech in a different setting. This presents a compelling alternative tuning strategy for speech recognition using large language models.
title Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2501.00039