Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Peng, Zhu, Yanqiao, Jiang, Zixuan, Chen, Qinyuan, Zhao, Xingjian, Qiu, Xipeng, Wang, Wupeng, Gao, Zhifu, Li, Xiangang, Yu, Kai, Chen, Xie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Explore the Reinforcement Learning for the LLM based ASR and TTS system
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
RAS: a Reliability Oriented Metric for Automatic Speech Recognition
von: Huang, Wenbin, et al.
Veröffentlicht: (2026)
von: Huang, Wenbin, et al.
Veröffentlicht: (2026)
Fun-ASR Technical Report
von: An, Keyu, et al.
Veröffentlicht: (2025)
von: An, Keyu, et al.
Veröffentlicht: (2025)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
von: Yang, Chenchen, et al.
Veröffentlicht: (2026)
von: Yang, Chenchen, et al.
Veröffentlicht: (2026)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
von: Wang, He, et al.
Veröffentlicht: (2025)
von: Wang, He, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models
von: Feng, Chen, et al.
Veröffentlicht: (2025)
von: Feng, Chen, et al.
Veröffentlicht: (2025)
SpeechAlign: Aligning Speech Generation to Human Preferences
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
von: Deng, Ruifan, et al.
Veröffentlicht: (2025)
von: Deng, Ruifan, et al.
Veröffentlicht: (2025)
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
von: Liu, Zhanxun, et al.
Veröffentlicht: (2025)
von: Liu, Zhanxun, et al.
Veröffentlicht: (2025)
ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System
von: Xu, Kaituo, et al.
Veröffentlicht: (2026)
von: Xu, Kaituo, et al.
Veröffentlicht: (2026)
Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
UniVocal: Unified Speech-Singing Code-Switching Synthesis
von: Shi, Yufei, et al.
Veröffentlicht: (2026)
von: Shi, Yufei, et al.
Veröffentlicht: (2026)
When Tone and Words Disagree: Towards Robust Speech Emotion Recognition under Acoustic-Semantic Conflict
von: Huang, Dawei, et al.
Veröffentlicht: (2026)
von: Huang, Dawei, et al.
Veröffentlicht: (2026)
FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
von: Xu, Kai-Tuo, et al.
Veröffentlicht: (2025)
von: Xu, Kai-Tuo, et al.
Veröffentlicht: (2025)
Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
von: Wu, Weijie, et al.
Veröffentlicht: (2025)
von: Wu, Weijie, et al.
Veröffentlicht: (2025)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
von: Wang, Wupeng, et al.
Veröffentlicht: (2024)
von: Wang, Wupeng, et al.
Veröffentlicht: (2024)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Towards Decoupling Frontend Enhancement and Backend Recognition in Monaural Robust ASR
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
von: Wei, Linye, et al.
Veröffentlicht: (2025)
von: Wei, Linye, et al.
Veröffentlicht: (2025)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
CLiFT-ASR: A Cross-Lingual Fine-Tuning Framework for Low-Resource Taiwanese Hokkien Speech Recognition
von: Sung, Hung-Yang, et al.
Veröffentlicht: (2025)
von: Sung, Hung-Yang, et al.
Veröffentlicht: (2025)
Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2026)
von: Chen, Youjun, et al.
Veröffentlicht: (2026)
DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
von: Liu, Qingyu, et al.
Veröffentlicht: (2025)
von: Liu, Qingyu, et al.
Veröffentlicht: (2025)
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
CLAR: CIF-Localized Alignment for Retrieval-Augmented Speech LLM-Based Contextual ASR
von: Huang, Shangkun, et al.
Veröffentlicht: (2026)
von: Huang, Shangkun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Explore the Reinforcement Learning for the LLM based ASR and TTS system
von: Gao, Changfeng, et al.
Veröffentlicht: (2025) -
RAS: a Reliability Oriented Metric for Automatic Speech Recognition
von: Huang, Wenbin, et al.
Veröffentlicht: (2026) -
Fun-ASR Technical Report
von: An, Keyu, et al.
Veröffentlicht: (2025) -
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024) -
ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)