Attention-guided Evidence Grounding for Spoken Question Answering

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Ke, Chen, Bolin, Li, Yuejie, Hua, Yueying, Nie, Jianhao, He, Yueping, Li, Bowen, Mao, Chengjun
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917350946635776
author Yang, Ke
Chen, Bolin
Li, Yuejie
Hua, Yueying
Nie, Jianhao
He, Yueping
Li, Bowen
Mao, Chengjun
author_facet Yang, Ke
Chen, Bolin
Li, Yuejie
Hua, Yueying
Nie, Jianhao
He, Yueping
Li, Bowen
Mao, Chengjun
contents Spoken Question Answering (Spoken QA) presents a challenging cross-modal problem: effectively aligning acoustic queries with textual knowledge while avoiding the latency and error propagation inherent in cascaded ASR-based systems. In this paper, we introduce Attention-guided Evidence Grounding (AEG), a novel end-to-end framework that leverages the internal cross-modal attention of Speech Large Language Models (SpeechLLMs) to explicitly locate and ground key evidence in the model's latent space. To address the diffuse attention distribution in pre-trained models, we propose Learning to Focus on Evidence (LFE), a supervised fine-tuning paradigm that calibrates the model's attention mechanism to distinguish query-relevant segments from irrelevant context. Experiments on SQuAD, HotpotQA, and MuSiQue demonstrate that AEG reduces hallucinations and achieves strong efficiency gains, outperforming large-scale cascaded baselines (Whisper-Large-v3 + Reranker) while reducing inference latency by approximately 62%.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16292
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Attention-guided Evidence Grounding for Spoken Question Answering
Yang, Ke
Chen, Bolin
Li, Yuejie
Hua, Yueying
Nie, Jianhao
He, Yueping
Li, Bowen
Mao, Chengjun
Computation and Language
Artificial Intelligence
Spoken Question Answering (Spoken QA) presents a challenging cross-modal problem: effectively aligning acoustic queries with textual knowledge while avoiding the latency and error propagation inherent in cascaded ASR-based systems. In this paper, we introduce Attention-guided Evidence Grounding (AEG), a novel end-to-end framework that leverages the internal cross-modal attention of Speech Large Language Models (SpeechLLMs) to explicitly locate and ground key evidence in the model's latent space. To address the diffuse attention distribution in pre-trained models, we propose Learning to Focus on Evidence (LFE), a supervised fine-tuning paradigm that calibrates the model's attention mechanism to distinguish query-relevant segments from irrelevant context. Experiments on SQuAD, HotpotQA, and MuSiQue demonstrate that AEG reduces hallucinations and achieves strong efficiency gains, outperforming large-scale cascaded baselines (Whisper-Large-v3 + Reranker) while reducing inference latency by approximately 62%.
title Attention-guided Evidence Grounding for Spoken Question Answering
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.16292