LaSR: Context-Aware Speech Recognition via Latent Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Heyang, Cheng, Ziyang, Huang, Jiayi, Xiao, Wenyang, Wu, Ronghua, Gu, Qunshan, Wang, Yanfeng, Wang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
by: Cheng, Ziyang, et al.
Published: (2026)
by: Cheng, Ziyang, et al.
Published: (2026)
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
by: Hou, Yixuan, et al.
Published: (2025)
by: Hou, Yixuan, et al.
Published: (2025)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
by: Liu, Heyang, et al.
Published: (2024)
by: Liu, Heyang, et al.
Published: (2024)
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition
by: Rong, Yiming, et al.
Published: (2025)
by: Rong, Yiming, et al.
Published: (2025)
Towards an End-to-End Framework for Invasive Brain Signal Decoding with Large Language Models
by: Feng, Sheng, et al.
Published: (2024)
by: Feng, Sheng, et al.
Published: (2024)
Context-Aware Dynamic Chunking for Streaming Tibetan Speech Recognition
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models
by: Zhao, Zihan, et al.
Published: (2023)
by: Zhao, Zihan, et al.
Published: (2023)
WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
The Role of Diversity in In-Context Learning for Large Language Models
by: Xiao, Wenyang, et al.
Published: (2025)
by: Xiao, Wenyang, et al.
Published: (2025)
Latent Thoughts Tuning: Bridging Context and Reasoning with Fused Information in Latent Tokens
by: Liu, Weihao, et al.
Published: (2026)
by: Liu, Weihao, et al.
Published: (2026)
Pipelined Decoder for Efficient Context-Aware Text Generation
by: Huang, Zixian, et al.
Published: (2025)
by: Huang, Zixian, et al.
Published: (2025)
Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm
by: Liu, Hongcheng, et al.
Published: (2024)
by: Liu, Hongcheng, et al.
Published: (2024)
M$^3$AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset
by: Chen, Zhe, et al.
Published: (2024)
by: Chen, Zhe, et al.
Published: (2024)
Bridging the Language Gap: Synthetic Voice Diversity via Latent Mixup for Equitable Speech Recognition
by: Bian, Wesley, et al.
Published: (2025)
by: Bian, Wesley, et al.
Published: (2025)
Leveraging Diverse Modeling Contexts with Collaborating Learning for Neural Machine Translation
by: Liao, Yusheng, et al.
Published: (2024)
by: Liao, Yusheng, et al.
Published: (2024)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
by: Petty, Jackson, et al.
Published: (2025)
by: Petty, Jackson, et al.
Published: (2025)
Long Context Scaling: Divide and Conquer via Multi-Agent Question-driven Collaboration
by: Xiao, Sibo, et al.
Published: (2025)
by: Xiao, Sibo, et al.
Published: (2025)
An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition
by: Wang, Yi-Cheng, et al.
Published: (2024)
by: Wang, Yi-Cheng, et al.
Published: (2024)
Improving Context Fidelity via Native Retrieval-Augmented Reasoning
by: Wang, Suyuchen, et al.
Published: (2025)
by: Wang, Suyuchen, et al.
Published: (2025)
Context Pruning for Coding Agents via Multi-Rubric Latent Reasoning
by: Wang, Jingjing, et al.
Published: (2026)
by: Wang, Jingjing, et al.
Published: (2026)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
by: Gui, Jiayi, et al.
Published: (2024)
by: Gui, Jiayi, et al.
Published: (2024)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
by: Wang, Yujin, et al.
Published: (2022)
by: Wang, Yujin, et al.
Published: (2022)
ComPASS: Towards Personalized Agentic Social Support via Tool-Augmented Companionship
by: Huang, Zhaopei, et al.
Published: (2026)
by: Huang, Zhaopei, et al.
Published: (2026)
ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents
by: Liao, Yusheng, et al.
Published: (2024)
by: Liao, Yusheng, et al.
Published: (2024)
LaTER: Efficient Test-Time Reasoning via Latent Exploration and Explicit Verification
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
Enhancing Long-Chain Reasoning Distillation through Error-Aware Self-Reflection
by: Wu, Zhuoyang, et al.
Published: (2025)
by: Wu, Zhuoyang, et al.
Published: (2025)
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
by: Cha, Jun-Hyeok, et al.
Published: (2025)
by: Cha, Jun-Hyeok, et al.
Published: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
by: Lin, Hsi-Che, et al.
Published: (2024)
by: Lin, Hsi-Che, et al.
Published: (2024)
SR-KI: Scalable and Real-Time Knowledge Integration into LLMs via Supervised Attention
by: Yu, Bohan, et al.
Published: (2025)
by: Yu, Bohan, et al.
Published: (2025)
Latent Visual Reasoning
by: Li, Bangzheng, et al.
Published: (2025)
by: Li, Bangzheng, et al.
Published: (2025)
TARQ: Tail-Aware Reconstruction Quantization for Rare-Word Robust Automatic Speech Recognition
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Similar Items
-
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
by: Liu, Heyang, et al.
Published: (2025) -
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
by: Cheng, Ziyang, et al.
Published: (2026) -
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
by: Liu, Heyang, et al.
Published: (2025) -
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
by: Wang, Yuhao, et al.
Published: (2025) -
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
by: Wang, Yuhao, et al.
Published: (2025)