Head Pursuit: Probing Attention Specialization in Multimodal Transformers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Basile, Lorenzo, Maiorca, Valentino, Doimo, Diego, Locatello, Francesco, Cazzaniga, Alberto
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914255375171584
author Basile, Lorenzo
Maiorca, Valentino
Doimo, Diego
Locatello, Francesco
Cazzaniga, Alberto
author_facet Basile, Lorenzo
Maiorca, Valentino
Doimo, Diego
Locatello, Francesco
Cazzaniga, Alberto
contents Language and vision-language models have shown impressive performance across a wide range of tasks, but their internal mechanisms remain only partly understood. In this work, we study how individual attention heads in text-generative models specialize in specific semantic or visual attributes. Building on an established interpretability method, we reinterpret the practice of probing intermediate activations with the final decoding layer through the lens of signal processing. This lets us analyze multiple samples in a principled way and rank attention heads based on their relevance to target concepts. Our results show consistent patterns of specialization at the head level across both unimodal and multimodal transformers. Remarkably, we find that editing as few as 1% of the heads, selected using our method, can reliably suppress or enhance targeted concepts in the model output. We validate our approach on language tasks such as question answering and toxicity mitigation, as well as vision-language tasks including image classification and captioning. Our findings highlight an interpretable and controllable structure within attention layers, offering simple tools for understanding and editing large-scale generative models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21518
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Head Pursuit: Probing Attention Specialization in Multimodal Transformers
Basile, Lorenzo
Maiorca, Valentino
Doimo, Diego
Locatello, Francesco
Cazzaniga, Alberto
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Language and vision-language models have shown impressive performance across a wide range of tasks, but their internal mechanisms remain only partly understood. In this work, we study how individual attention heads in text-generative models specialize in specific semantic or visual attributes. Building on an established interpretability method, we reinterpret the practice of probing intermediate activations with the final decoding layer through the lens of signal processing. This lets us analyze multiple samples in a principled way and rank attention heads based on their relevance to target concepts. Our results show consistent patterns of specialization at the head level across both unimodal and multimodal transformers. Remarkably, we find that editing as few as 1% of the heads, selected using our method, can reliably suppress or enhance targeted concepts in the model output. We validate our approach on language tasks such as question answering and toxicity mitigation, as well as vision-language tasks including image classification and captioning. Our findings highlight an interpretable and controllable structure within attention layers, offering simple tools for understanding and editing large-scale generative models.
title Head Pursuit: Probing Attention Specialization in Multimodal Transformers
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.21518