Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anand, Cappellazzo, Umberto, Petridis, Stavros, Pantic, Maja
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911402822729728
author Anand
Cappellazzo, Umberto
Petridis, Stavros
Pantic, Maja
author_facet Anand
Cappellazzo, Umberto
Petridis, Stavros
Pantic, Maja
contents Large language models (LLMs) have recently advanced auditory speech recognition (ASR), visual speech recognition (VSR), and audio-visual speech recognition (AVSR). However, understanding of their internal dynamics under fine-tuning remains limited. In natural language processing, recent work has revealed attention sinks, tokens that attract disproportionately high attention, and associated massive activations in which some features of sink tokens exhibit huge activation in LLMs. In this work, we are the first to study these phenomena in multimodal speech recognition. Through a detailed analysis of audio-visual LLMs, we identify attention sinks and massive activations not only at the BOS token but also at intermediate low-semantic tokens across ASR, VSR, and AVSR. We show that massive activations originate in the MLP layers and correspond to fixed feature indices across all sink tokens. We further show that intermediate sink tokens exhibit high cosine similarity to the BOS token, thereby amplifying attention and activation. Building on these insights, we introduce a simple decorrelation loss that reduces cosine similarity between BOS and other tokens, effectively mitigating intermediate sinks and massive activations. Furthermore, our method improves word error rate (WER) under high audio-visual feature downsampling while remaining stable at lower downsampling rates.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22603
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
Anand
Cappellazzo, Umberto
Petridis, Stavros
Pantic, Maja
Audio and Speech Processing
Computer Vision and Pattern Recognition
Sound
Large language models (LLMs) have recently advanced auditory speech recognition (ASR), visual speech recognition (VSR), and audio-visual speech recognition (AVSR). However, understanding of their internal dynamics under fine-tuning remains limited. In natural language processing, recent work has revealed attention sinks, tokens that attract disproportionately high attention, and associated massive activations in which some features of sink tokens exhibit huge activation in LLMs. In this work, we are the first to study these phenomena in multimodal speech recognition. Through a detailed analysis of audio-visual LLMs, we identify attention sinks and massive activations not only at the BOS token but also at intermediate low-semantic tokens across ASR, VSR, and AVSR. We show that massive activations originate in the MLP layers and correspond to fixed feature indices across all sink tokens. We further show that intermediate sink tokens exhibit high cosine similarity to the BOS token, thereby amplifying attention and activation. Building on these insights, we introduce a simple decorrelation loss that reduces cosine similarity between BOS and other tokens, effectively mitigating intermediate sinks and massive activations. Furthermore, our method improves word error rate (WER) under high audio-visual feature downsampling while remaining stable at lower downsampling rates.
title Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
topic Audio and Speech Processing
Computer Vision and Pattern Recognition
Sound
url https://arxiv.org/abs/2510.22603