Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
Fuente:
arXiv
Guardado en:
| Autores principales: | Lin, Feng, Chen, Marco, Zhang, Haokui, Yu, Xiaotian, Lu, Guangming, Xiao, Rong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Attention Is All You Need For Mixture-of-Depths Routing
por: Gadhikar, Advait, et al.
Publicado: (2024)
por: Gadhikar, Advait, et al.
Publicado: (2024)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
por: Yu, Lu, et al.
Publicado: (2024)
por: Yu, Lu, et al.
Publicado: (2024)
Transferable-guided Attention Is All You Need for Video Domain Adaptation
por: Sacilotti, André, et al.
Publicado: (2024)
por: Sacilotti, André, et al.
Publicado: (2024)
CLIP is All You Need for Human-like Semantic Representations in Stable Diffusion
por: Braunstein, Cameron, et al.
Publicado: (2025)
por: Braunstein, Cameron, et al.
Publicado: (2025)
Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps
por: Kim, Jeeyung, et al.
Publicado: (2024)
por: Kim, Jeeyung, et al.
Publicado: (2024)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
por: Zhao, Kairan, et al.
Publicado: (2026)
por: Zhao, Kairan, et al.
Publicado: (2026)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
por: Song, Junha, et al.
Publicado: (2026)
por: Song, Junha, et al.
Publicado: (2026)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
por: Yeo, Wei Jie, et al.
Publicado: (2025)
por: Yeo, Wei Jie, et al.
Publicado: (2025)
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
por: Lan, Mengcheng, et al.
Publicado: (2024)
por: Lan, Mengcheng, et al.
Publicado: (2024)
You Only Need Less Attention at Each Stage in Vision Transformers
por: Zhang, Shuoxi, et al.
Publicado: (2024)
por: Zhang, Shuoxi, et al.
Publicado: (2024)
Bridging Sensor Gaps via Attention Gated Tuning for Hyperspectral Image Classification
por: Xue, Xizhe, et al.
Publicado: (2023)
por: Xue, Xizhe, et al.
Publicado: (2023)
Memory augment is All You Need for image restoration
por: Zhang, Xiao Feng, et al.
Publicado: (2023)
por: Zhang, Xiao Feng, et al.
Publicado: (2023)
Attention Head Purification: A New Perspective to Harness CLIP for Domain Generalization
por: Wang, Yingfan, et al.
Publicado: (2024)
por: Wang, Yingfan, et al.
Publicado: (2024)
Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
por: Bui, Phuoc-Nguyen, et al.
Publicado: (2025)
por: Bui, Phuoc-Nguyen, et al.
Publicado: (2025)
Generalize Your Face Forgery Detectors: An Insertable Adaptation Module Is All You Need
por: Si, Xiaotian, et al.
Publicado: (2024)
por: Si, Xiaotian, et al.
Publicado: (2024)
TANet: Triplet Attention Network for All-In-One Adverse Weather Image Restoration
por: Wang, Hsing-Hua, et al.
Publicado: (2024)
por: Wang, Hsing-Hua, et al.
Publicado: (2024)
Multi-View Representation is What You Need for Point-Cloud Pre-Training
por: Yan, Siming, et al.
Publicado: (2023)
por: Yan, Siming, et al.
Publicado: (2023)
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
por: Tang, Lexiang, et al.
Publicado: (2025)
por: Tang, Lexiang, et al.
Publicado: (2025)
MaTe: Images Are All You Need for Material Transfer via Diffusion Transformer
por: Huang, Nisha, et al.
Publicado: (2026)
por: Huang, Nisha, et al.
Publicado: (2026)
Prompt-Guided Attention Head Selection for Focus-Oriented Image Retrieval
por: Nozawa, Yuji, et al.
Publicado: (2025)
por: Nozawa, Yuji, et al.
Publicado: (2025)
Spatio-Temporal Progressive Attention Model for EEG Classification in Rapid Serial Visual Presentation Task
por: Li, Yang, et al.
Publicado: (2025)
por: Li, Yang, et al.
Publicado: (2025)
Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
por: Cao, Pu, et al.
Publicado: (2023)
por: Cao, Pu, et al.
Publicado: (2023)
Locating Demographic Bias at the Attention-Head Level in CLIP's Vision Encoder
por: Yasser, Alaa, et al.
Publicado: (2026)
por: Yasser, Alaa, et al.
Publicado: (2026)
A Single Simple Patch is All You Need for AI-generated Image Detection
por: Chen, Jiaxuan, et al.
Publicado: (2024)
por: Chen, Jiaxuan, et al.
Publicado: (2024)
Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads
por: Rahimian, Ali K., et al.
Publicado: (2024)
por: Rahimian, Ali K., et al.
Publicado: (2024)
Pairwise Comparisons Are All You Need
por: Chahine, Nicolas, et al.
Publicado: (2024)
por: Chahine, Nicolas, et al.
Publicado: (2024)
Similarity Memory Prior is All You Need for Medical Image Segmentation
por: Tang, Hao, et al.
Publicado: (2025)
por: Tang, Hao, et al.
Publicado: (2025)
All You Need to Know About Training Image Retrieval Models
por: Berton, Gabriele, et al.
Publicado: (2025)
por: Berton, Gabriele, et al.
Publicado: (2025)
Location Is All You Need: Continuous Spatiotemporal Neural Representations of Earth Observation Data
por: Madadikhaljan, Mojgan, et al.
Publicado: (2026)
por: Madadikhaljan, Mojgan, et al.
Publicado: (2026)
Smart Feature is What You Need
por: Hu, Zhaoxin, et al.
Publicado: (2024)
por: Hu, Zhaoxin, et al.
Publicado: (2024)
COCO is "ALL'' You Need for Visual Instruction Fine-tuning
por: Han, Xiaotian, et al.
Publicado: (2024)
por: Han, Xiaotian, et al.
Publicado: (2024)
A Single Image and Multimodality Is All You Need for Novel View Synthesis
por: Javadi, Amirhosein, et al.
Publicado: (2026)
por: Javadi, Amirhosein, et al.
Publicado: (2026)
ParameterNet: Parameters Are All You Need
por: Han, Kai, et al.
Publicado: (2023)
por: Han, Kai, et al.
Publicado: (2023)
[MASK] is All You Need
por: Hu, Vincent Tao, et al.
Publicado: (2024)
por: Hu, Vincent Tao, et al.
Publicado: (2024)
SMPL Normal Map Is All You Need for Single-view Textured Human Reconstruction
por: Shen, Wenhao, et al.
Publicado: (2025)
por: Shen, Wenhao, et al.
Publicado: (2025)
CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
por: Lavoie, Marc-Antoine, et al.
Publicado: (2026)
por: Lavoie, Marc-Antoine, et al.
Publicado: (2026)
Emu3: Next-Token Prediction is All You Need
por: Wang, Xinlong, et al.
Publicado: (2024)
por: Wang, Xinlong, et al.
Publicado: (2024)
Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
por: Li, Zhiqi, et al.
Publicado: (2023)
por: Li, Zhiqi, et al.
Publicado: (2023)
Anatomy Might Be All You Need: Forecasting What to Do During Surgery
por: Sarwin, Gary, et al.
Publicado: (2025)
por: Sarwin, Gary, et al.
Publicado: (2025)
DiffCLIP: Differential Attention Meets CLIP
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2025)
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2025)
Ejemplares similares
-
Attention Is All You Need For Mixture-of-Depths Routing
por: Gadhikar, Advait, et al.
Publicado: (2024) -
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
por: Yu, Lu, et al.
Publicado: (2024) -
Transferable-guided Attention Is All You Need for Video Domain Adaptation
por: Sacilotti, André, et al.
Publicado: (2024) -
CLIP is All You Need for Human-like Semantic Representations in Stable Diffusion
por: Braunstein, Cameron, et al.
Publicado: (2025) -
Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps
por: Kim, Jeeyung, et al.
Publicado: (2024)