Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Feng, Chen, Marco, Zhang, Haokui, Yu, Xiaotian, Lu, Guangming, Xiao, Rong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914159520645120
author Lin, Feng
Chen, Marco
Zhang, Haokui
Yu, Xiaotian
Lu, Guangming
Xiao, Rong
author_facet Lin, Feng
Chen, Marco
Zhang, Haokui
Yu, Xiaotian
Lu, Guangming
Xiao, Rong
contents This paper investigates the role of attention heads in CLIP's image encoder. Building on interpretability studies, we conduct an exhaustive analysis and find that certain heads, distributed across layers, are detrimental to the resulting representations. To mitigate their impact, we propose a simple yet effective Attention Ablation Technique (AAT) that suppresses selected heads by directly manipulating their attention weights. By incorporating two complementary strategies tailored to different application scenarios, AAT enables the systematic identification and ablation of harmful heads with minimal overhead. Experiments show that AAT consistently improves downstream performance across diverse domains, boosting recall by up to 11.1% on cross-modal retrieval benchmarks. These results highlight that AAT can effectively refine large-scale VLMs with virtually no extra inference cost, while yielding semantically meaningful patterns that align with existing interpretability findings.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00537
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
Lin, Feng
Chen, Marco
Zhang, Haokui
Yu, Xiaotian
Lu, Guangming
Xiao, Rong
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
This paper investigates the role of attention heads in CLIP's image encoder. Building on interpretability studies, we conduct an exhaustive analysis and find that certain heads, distributed across layers, are detrimental to the resulting representations. To mitigate their impact, we propose a simple yet effective Attention Ablation Technique (AAT) that suppresses selected heads by directly manipulating their attention weights. By incorporating two complementary strategies tailored to different application scenarios, AAT enables the systematic identification and ablation of harmful heads with minimal overhead. Experiments show that AAT consistently improves downstream performance across diverse domains, boosting recall by up to 11.1% on cross-modal retrieval benchmarks. These results highlight that AAT can effectively refine large-scale VLMs with virtually no extra inference cost, while yielding semantically meaningful patterns that align with existing interpretability findings.
title Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.00537