Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yiru, Su, Hang, Fan, Lichun, Luo, Zhenbo, Luan, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912860706177024
author Zhang, Yiru
Su, Hang
Fan, Lichun
Luo, Zhenbo
Luan, Jian
author_facet Zhang, Yiru
Su, Hang
Fan, Lichun
Luo, Zhenbo
Luan, Jian
contents Target Speaker Automatic Speech Recognition (TS-ASR) aims to transcribe the speech of a specified target speaker from multi-speaker mixtures in cocktail party scenarios. Recent advancement of Large Audio-Language Models (LALMs) has already brought some new insights to TS-ASR. However, significant room for optimization remains for the TS-ASR task within the LALMs architecture. While Chain of Thoughts (CoT) and Reinforcement Learning (RL) have proven effective in certain speech tasks, TS-ASR, which requires the model to deeply comprehend speech signals, differentiate various speakers, and handle overlapping utterances is particularly well-suited to a reasoning-guided approach. Therefore, we propose a novel framework that incorporates CoT and RL training into TS-ASR for performance improvement. A novel CoT dataset of TS-ASR is constructed, and the TS-ASR model is first trained on regular data and then fine-tuned on CoT data. Finally, the model is further trained with RL using selected data to enhance generalized reasoning capabilities. Experiment results show a significant improvement of TS-ASR performance with CoT and RL training, which demonstrates the effectiveness of the proposed CoT and RL training methods adapted for the TS-ASR task.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15612
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
Zhang, Yiru
Su, Hang
Fan, Lichun
Luo, Zhenbo
Luan, Jian
Sound
Audio and Speech Processing
Target Speaker Automatic Speech Recognition (TS-ASR) aims to transcribe the speech of a specified target speaker from multi-speaker mixtures in cocktail party scenarios. Recent advancement of Large Audio-Language Models (LALMs) has already brought some new insights to TS-ASR. However, significant room for optimization remains for the TS-ASR task within the LALMs architecture. While Chain of Thoughts (CoT) and Reinforcement Learning (RL) have proven effective in certain speech tasks, TS-ASR, which requires the model to deeply comprehend speech signals, differentiate various speakers, and handle overlapping utterances is particularly well-suited to a reasoning-guided approach. Therefore, we propose a novel framework that incorporates CoT and RL training into TS-ASR for performance improvement. A novel CoT dataset of TS-ASR is constructed, and the TS-ASR model is first trained on regular data and then fine-tuned on CoT data. Finally, the model is further trained with RL using selected data to enhance generalized reasoning capabilities. Experiment results show a significant improvement of TS-ASR performance with CoT and RL training, which demonstrates the effectiveness of the proposed CoT and RL training methods adapted for the TS-ASR task.
title Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.15612