Explore the Reinforcement Learning for the LLM based ASR and TTS system

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Changfeng, Li, Yabin, An, Keyu, Gao, Zhifu, Du, Zhihao, Zhao, Han, Li, Xiangang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908554822156288
author Gao, Changfeng
Li, Yabin
An, Keyu
Gao, Zhifu
Du, Zhihao
Zhao, Han
Li, Xiangang
author_facet Gao, Changfeng
Li, Yabin
An, Keyu
Gao, Zhifu
Du, Zhihao
Zhao, Han
Li, Xiangang
contents In recent years, large language models (LLMs) have played an important role in automatic speech recognition (ASR) and text-to-speech (TTS) systems. While reinforcement learning (RL) has significantly enhanced LLM performance in text-based tasks, its application to ASR and TTS remains underexplored due to the complexity of training audio-based models. In this study, we propose a lightweight RL framework tailored for audio-based LLMs that can process audio inputs and generate audio outputs. Based on this framework, we evaluate the effectiveness of reinforcement learning on both ASR and TTS tasks. For the ASR task, we experiment with different rule-based reward functions within the Group Relative Policy Optimization (GRPO) framework and investigate the impact of RL data construction. For the TTS task, we compare GRPO with Differentiable Reward Optimization (DiffRO) and further combine the two approaches to achieve improved performance. Our experiments demonstrate that RL can significantly enhance the performance of both ASR and TTS systems, even with limited training data and a small number of optimization steps.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18569
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Explore the Reinforcement Learning for the LLM based ASR and TTS system
Gao, Changfeng
Li, Yabin
An, Keyu
Gao, Zhifu
Du, Zhihao
Zhao, Han
Li, Xiangang
Sound
Artificial Intelligence
Audio and Speech Processing
In recent years, large language models (LLMs) have played an important role in automatic speech recognition (ASR) and text-to-speech (TTS) systems. While reinforcement learning (RL) has significantly enhanced LLM performance in text-based tasks, its application to ASR and TTS remains underexplored due to the complexity of training audio-based models. In this study, we propose a lightweight RL framework tailored for audio-based LLMs that can process audio inputs and generate audio outputs. Based on this framework, we evaluate the effectiveness of reinforcement learning on both ASR and TTS tasks. For the ASR task, we experiment with different rule-based reward functions within the Group Relative Policy Optimization (GRPO) framework and investigate the impact of RL data construction. For the TTS task, we compare GRPO with Differentiable Reward Optimization (DiffRO) and further combine the two approaches to achieve improved performance. Our experiments demonstrate that RL can significantly enhance the performance of both ASR and TTS systems, even with limited training data and a small number of optimization steps.
title Explore the Reinforcement Learning for the LLM based ASR and TTS system
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2509.18569