MoH: Multi-Head Attention as Mixture-of-Head Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Peng, Zhu, Bo, Yuan, Li, Yan, Shuicheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Image Generation with Variadic Attention Heads
by: Walton, Steven, et al.
Published: (2022)
by: Walton, Steven, et al.
Published: (2022)
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
by: Lin, Feng, et al.
Published: (2025)
by: Lin, Feng, et al.
Published: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Multi-Head Encoding for Extreme Label Classification
by: Liang, Daojun, et al.
Published: (2024)
by: Liang, Daojun, et al.
Published: (2024)
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
by: Constantinou, Christos, et al.
Published: (2024)
by: Constantinou, Christos, et al.
Published: (2024)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
by: Choi, Kanghyun, et al.
Published: (2024)
by: Choi, Kanghyun, et al.
Published: (2024)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
by: Fan, Xiaoran, et al.
Published: (2026)
by: Fan, Xiaoran, et al.
Published: (2026)
Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models
by: Zheng, Ziwei, et al.
Published: (2025)
by: Zheng, Ziwei, et al.
Published: (2025)
Automatic Channel Pruning for Multi-Head Attention
by: Lee, Eunho, et al.
Published: (2024)
by: Lee, Eunho, et al.
Published: (2024)
Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
by: Hassani, Ali, et al.
Published: (2025)
by: Hassani, Ali, et al.
Published: (2025)
MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
by: Liu, YiTong, et al.
Published: (2025)
by: Liu, YiTong, et al.
Published: (2025)
Multi-layer Learnable Attention Mask for Multimodal Tasks
by: Barrios, Wayner, et al.
Published: (2024)
by: Barrios, Wayner, et al.
Published: (2024)
COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision Models
by: Xiao, Jinqi, et al.
Published: (2023)
by: Xiao, Jinqi, et al.
Published: (2023)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers
by: Zhang, Hanling, et al.
Published: (2025)
by: Zhang, Hanling, et al.
Published: (2025)
Decoupling Amplitude and Phase Attention in Frequency Domain for RGB-Event based Visual Object Tracking
by: Wang, Shiao, et al.
Published: (2026)
by: Wang, Shiao, et al.
Published: (2026)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
by: Li, Xingyang, et al.
Published: (2025)
by: Li, Xingyang, et al.
Published: (2025)
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
by: Xu, Lu, et al.
Published: (2025)
by: Xu, Lu, et al.
Published: (2025)
Object-level Cross-view Geo-localization with Location Enhancement and Multi-Head Cross Attention
by: Huang, Zheyang, et al.
Published: (2025)
by: Huang, Zheyang, et al.
Published: (2025)
HydraViT: Stacking Heads for a Scalable ViT
by: Haberer, Janek, et al.
Published: (2024)
by: Haberer, Janek, et al.
Published: (2024)
Fairness-aware Vision Transformer via Debiased Self-Attention
by: Qiang, Yao, et al.
Published: (2023)
by: Qiang, Yao, et al.
Published: (2023)
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
by: Zhang, Kewei, et al.
Published: (2026)
by: Zhang, Kewei, et al.
Published: (2026)
I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts
by: Xin, Jiayi, et al.
Published: (2025)
by: Xin, Jiayi, et al.
Published: (2025)
Exploring the Stability Gap in Continual Learning: The Role of the Classification Head
by: Łapacz, Wojciech, et al.
Published: (2024)
by: Łapacz, Wojciech, et al.
Published: (2024)
InceptionNeXt: When Inception Meets ConvNeXt
by: Yu, Weihao, et al.
Published: (2023)
by: Yu, Weihao, et al.
Published: (2023)
Precipitation Nowcasting Using Diffusion Transformer with Causal Attention
by: Li, ChaoRong, et al.
Published: (2024)
by: Li, ChaoRong, et al.
Published: (2024)
Masked Multi-Query Slot Attention for Unsupervised Object Discovery
by: Pramanik, Rishav, et al.
Published: (2024)
by: Pramanik, Rishav, et al.
Published: (2024)
Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation
by: Xu, Jie, et al.
Published: (2025)
by: Xu, Jie, et al.
Published: (2025)
Attention in Diffusion Model: A Survey
by: Hua, Litao, et al.
Published: (2025)
by: Hua, Litao, et al.
Published: (2025)
Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks
by: Łucki, Jakub, et al.
Published: (2025)
by: Łucki, Jakub, et al.
Published: (2025)
Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
by: Colagrande, Alex, et al.
Published: (2025)
by: Colagrande, Alex, et al.
Published: (2025)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026)
by: Ma, Xueqi, et al.
Published: (2026)
OTSeg: Multi-prompt Sinkhorn Attention for Zero-Shot Semantic Segmentation
by: Kim, Kwanyoung, et al.
Published: (2024)
by: Kim, Kwanyoung, et al.
Published: (2024)
GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers
by: Miyato, Takeru, et al.
Published: (2023)
by: Miyato, Takeru, et al.
Published: (2023)
Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots
by: Lao, Dong, et al.
Published: (2023)
by: Lao, Dong, et al.
Published: (2023)
Attention-ResUNet for Automated Fetal Head Segmentation
by: Bhilwarawala, Ammar, et al.
Published: (2026)
by: Bhilwarawala, Ammar, et al.
Published: (2026)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts
by: Zhu, Jie, et al.
Published: (2024)
by: Zhu, Jie, et al.
Published: (2024)
Similar Items
-
Efficient Image Generation with Variadic Attention Heads
by: Walton, Steven, et al.
Published: (2022) -
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
by: Lin, Feng, et al.
Published: (2025) -
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025) -
Multi-Head Encoding for Extreme Label Classification
by: Liang, Daojun, et al.
Published: (2024) -
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
by: Constantinou, Christos, et al.
Published: (2024)