Cross-Attention is Half Explanation in Speech-to-Text Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Papi, Sara, Fucci, Dennis, Gaido, Marco, Negri, Matteo, Bentivogli, Luisa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914051216375808
author Papi, Sara
Fucci, Dennis
Gaido, Marco
Negri, Matteo
Bentivogli, Luisa
author_facet Papi, Sara
Fucci, Dennis
Gaido, Marco
Negri, Matteo
Bentivogli, Luisa
contents Cross-attention is a core mechanism in encoder-decoder architectures, widespread in many fields, including speech-to-text (S2T) processing. Its scores have been repurposed for various downstream applications--such as timestamp estimation and audio-text alignment--under the assumption that they reflect the dependencies between input speech representation and the generated text. While the explanatory nature of attention mechanisms has been widely debated in the broader NLP literature, this assumption remains largely unexplored within the speech domain. To address this gap, we assess the explanatory power of cross-attention in S2T models by comparing its scores to input saliency maps derived from feature attribution. Our analysis spans monolingual and multilingual, single-task and multi-task models at multiple scales, and shows that attention scores moderately to strongly align with saliency-based explanations, particularly when aggregated across heads and layers. However, it also shows that cross-attention captures only about 50% of the input relevance and, in the best case, only partially reflects how the decoder attends to the encoder's representations--accounting for just 52-75% of the saliency. These findings uncover fundamental limitations in interpreting cross-attention as an explanatory proxy, suggesting that it offers an informative yet incomplete view of the factors driving predictions in S2T models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18010
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cross-Attention is Half Explanation in Speech-to-Text Models
Papi, Sara
Fucci, Dennis
Gaido, Marco
Negri, Matteo
Bentivogli, Luisa
Computation and Language
Artificial Intelligence
Sound
Cross-attention is a core mechanism in encoder-decoder architectures, widespread in many fields, including speech-to-text (S2T) processing. Its scores have been repurposed for various downstream applications--such as timestamp estimation and audio-text alignment--under the assumption that they reflect the dependencies between input speech representation and the generated text. While the explanatory nature of attention mechanisms has been widely debated in the broader NLP literature, this assumption remains largely unexplored within the speech domain. To address this gap, we assess the explanatory power of cross-attention in S2T models by comparing its scores to input saliency maps derived from feature attribution. Our analysis spans monolingual and multilingual, single-task and multi-task models at multiple scales, and shows that attention scores moderately to strongly align with saliency-based explanations, particularly when aggregated across heads and layers. However, it also shows that cross-attention captures only about 50% of the input relevance and, in the best case, only partially reflects how the decoder attends to the encoder's representations--accounting for just 52-75% of the saliency. These findings uncover fundamental limitations in interpreting cross-attention as an explanatory proxy, suggesting that it offers an informative yet incomplete view of the factors driving predictions in S2T models.
title Cross-Attention is Half Explanation in Speech-to-Text Models
topic Computation and Language
Artificial Intelligence
Sound
url https://arxiv.org/abs/2509.18010