Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Peng, Zhu, Yanqiao, Jiang, Zixuan, Chen, Qinyuan, Zhao, Xingjian, Qiu, Xipeng, Wang, Wupeng, Gao, Zhifu, Li, Xiangang, Yu, Kai, Chen, Xie
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914470571278336
author Wang, Peng
Zhu, Yanqiao
Jiang, Zixuan
Chen, Qinyuan
Zhao, Xingjian
Qiu, Xipeng
Wang, Wupeng
Gao, Zhifu
Li, Xiangang
Yu, Kai
Chen, Xie
author_facet Wang, Peng
Zhu, Yanqiao
Jiang, Zixuan
Chen, Qinyuan
Zhao, Xingjian
Qiu, Xipeng
Wang, Wupeng
Gao, Zhifu
Li, Xiangang
Yu, Kai
Chen, Xie
contents Recent years have witnessed remarkable progress in automatic speech recognition (ASR), driven by advances in model architectures and large-scale training data. However, two important aspects remain underexplored. First, Word Error Rate (WER), the dominant evaluation metric for decades, treats all words equally and often fails to reflect the semantic correctness of an utterance at the sentence level. Second, interactive correction-an essential component of human communication-has rarely been systematically studied in ASR research. In this paper, we integrate these two perspectives under an agentic framework for interactive ASR. We propose leveraging LLM-as-a-Judge as a semantic-aware evaluation metric to assess recognition quality beyond token-level accuracy. Furthermore, we design an LLM-driven agent framework to simulate human-like multi-turn interaction, enabling iterative refinement of recognition outputs through semantic feedback. Extensive experiments are conducted on standard benchmarks, including GigaSpeech (English), WenetSpeech (Chinese), the ASRU 2019 code-switching test set. Both objective and subjective evaluations demonstrate the effectiveness of the proposed framework in improving semantic fidelity and interactive correction capability. We will release the code to facilitate future research in interactive and agentic ASR.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09121
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
Wang, Peng
Zhu, Yanqiao
Jiang, Zixuan
Chen, Qinyuan
Zhao, Xingjian
Qiu, Xipeng
Wang, Wupeng
Gao, Zhifu
Li, Xiangang
Yu, Kai
Chen, Xie
Computation and Language
Artificial Intelligence
Sound
Recent years have witnessed remarkable progress in automatic speech recognition (ASR), driven by advances in model architectures and large-scale training data. However, two important aspects remain underexplored. First, Word Error Rate (WER), the dominant evaluation metric for decades, treats all words equally and often fails to reflect the semantic correctness of an utterance at the sentence level. Second, interactive correction-an essential component of human communication-has rarely been systematically studied in ASR research. In this paper, we integrate these two perspectives under an agentic framework for interactive ASR. We propose leveraging LLM-as-a-Judge as a semantic-aware evaluation metric to assess recognition quality beyond token-level accuracy. Furthermore, we design an LLM-driven agent framework to simulate human-like multi-turn interaction, enabling iterative refinement of recognition outputs through semantic feedback. Extensive experiments are conducted on standard benchmarks, including GigaSpeech (English), WenetSpeech (Chinese), the ASRU 2019 code-switching test set. Both objective and subjective evaluations demonstrate the effectiveness of the proposed framework in improving semantic fidelity and interactive correction capability. We will release the code to facilitate future research in interactive and agentic ASR.
title Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
topic Computation and Language
Artificial Intelligence
Sound
url https://arxiv.org/abs/2604.09121