LaSR: Context-Aware Speech Recognition via Latent Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Heyang, Cheng, Ziyang, Huang, Jiayi, Xiao, Wenyang, Wu, Ronghua, Gu, Qunshan, Wang, Yanfeng, Wang, Yu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910273911128064
author Liu, Heyang
Cheng, Ziyang
Huang, Jiayi
Xiao, Wenyang
Wu, Ronghua
Gu, Qunshan
Wang, Yanfeng
Wang, Yu
author_facet Liu, Heyang
Cheng, Ziyang
Huang, Jiayi
Xiao, Wenyang
Wu, Ronghua
Gu, Qunshan
Wang, Yanfeng
Wang, Yu
contents Recent advances in Speech Large Language Models (Speech LLMs) have significantly enhanced spoken language understanding and reasoning. However, their contextual awareness is limited, struggling to perform speech recognition that effectively reflects the speaker's intent and topical context. In this paper, we propose LaSR (Latent Speech Reasoning), a novel training paradigm featuring a context-aware reasoning trajectory that leverages the latent reasoning process. Instead of generating explicit intermediate tokens, LaSR aligns chain-of-thought (CoT) supervision around the acoustic feature region of the targeted word, and introduces latent reasoning periods for context information grounding and transcriptional transition. Furthermore, to effectively benchmark contextual recognition on specialized vocabulary, we propose Spoken Darwin-Science, a large-scale corpus focusing on academic terminologies. Preliminary experiments on Fun-Audio-Chat demonstrate that LaSR significantly improves terminology recognition without introducing additional latency and consistently outperforms standard supervised fine-tuning baselines. Our findings highlight the potential of latent reasoning in building efficient, context-aware speech assistants.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00507
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LaSR: Context-Aware Speech Recognition via Latent Reasoning
Liu, Heyang
Cheng, Ziyang
Huang, Jiayi
Xiao, Wenyang
Wu, Ronghua
Gu, Qunshan
Wang, Yanfeng
Wang, Yu
Computation and Language
Recent advances in Speech Large Language Models (Speech LLMs) have significantly enhanced spoken language understanding and reasoning. However, their contextual awareness is limited, struggling to perform speech recognition that effectively reflects the speaker's intent and topical context. In this paper, we propose LaSR (Latent Speech Reasoning), a novel training paradigm featuring a context-aware reasoning trajectory that leverages the latent reasoning process. Instead of generating explicit intermediate tokens, LaSR aligns chain-of-thought (CoT) supervision around the acoustic feature region of the targeted word, and introduces latent reasoning periods for context information grounding and transcriptional transition. Furthermore, to effectively benchmark contextual recognition on specialized vocabulary, we propose Spoken Darwin-Science, a large-scale corpus focusing on academic terminologies. Preliminary experiments on Fun-Audio-Chat demonstrate that LaSR significantly improves terminology recognition without introducing additional latency and consistently outperforms standard supervised fine-tuning baselines. Our findings highlight the potential of latent reasoning in building efficient, context-aware speech assistants.
title LaSR: Context-Aware Speech Recognition via Latent Reasoning
topic Computation and Language
url https://arxiv.org/abs/2606.00507