Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers
Fuente:
arXiv
Saved in:
| Main Authors: | Bing, Zhaodong, Li, Linze, Liang, Jiajun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
by: Zhang, Shen, et al.
Published: (2025)
by: Zhang, Shen, et al.
Published: (2025)
MAT: Multi-Range Attention Transformer for Efficient Image Super-Resolution
by: Xie, Chengxing, et al.
Published: (2024)
by: Xie, Chengxing, et al.
Published: (2024)
Knowledge Distillation via the Target-aware Transformer
by: Lin, Sihao, et al.
Published: (2022)
by: Lin, Sihao, et al.
Published: (2022)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
MegActor-$Σ$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer
by: Yang, Shurong, et al.
Published: (2024)
by: Yang, Shurong, et al.
Published: (2024)
Frequency Attention for Knowledge Distillation
by: Pham, Cuong, et al.
Published: (2024)
by: Pham, Cuong, et al.
Published: (2024)
Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention into Convolutions
by: Dong, ZiYi, et al.
Published: (2025)
by: Dong, ZiYi, et al.
Published: (2025)
CrossKD: Cross-Head Knowledge Distillation for Object Detection
by: Wang, Jiabao, et al.
Published: (2023)
by: Wang, Jiabao, et al.
Published: (2023)
MoKD: Multi-Task Optimization for Knowledge Distillation
by: Hayder, Zeeshan, et al.
Published: (2025)
by: Hayder, Zeeshan, et al.
Published: (2025)
Asymmetric Decision-Making in Online Knowledge Distillation:Unifying Consensus and Divergence
by: Chen, Zhaowei, et al.
Published: (2025)
by: Chen, Zhaowei, et al.
Published: (2025)
The Velocity Deficit: Initial Energy Injection for Flow Matching
by: Li, Linze, et al.
Published: (2026)
by: Li, Linze, et al.
Published: (2026)
Efficient Temporal Sentence Grounding in Videos with Multi-Teacher Knowledge Distillation
by: Liang, Renjie, et al.
Published: (2023)
by: Liang, Renjie, et al.
Published: (2023)
Rethinking Centered Kernel Alignment in Knowledge Distillation
by: Zhou, Zikai, et al.
Published: (2024)
by: Zhou, Zikai, et al.
Published: (2024)
GMAR: Gradient-Driven Multi-Head Attention Rollout for Vision Transformer Interpretability
by: Jo, Sehyeong, et al.
Published: (2025)
by: Jo, Sehyeong, et al.
Published: (2025)
MHAFF: Multi-Head Attention Feature Fusion of CNN and Transformer for Cattle Identification
by: Dulal, Rabin, et al.
Published: (2025)
by: Dulal, Rabin, et al.
Published: (2025)
EVOKE: Emotion Enabled Virtual Avatar Mapping Using Optimized Knowledge Distillation
by: Nadeem, Maryam, et al.
Published: (2024)
by: Nadeem, Maryam, et al.
Published: (2024)
MegActor: Harness the Power of Raw Video for Vivid Portrait Animation
by: Yang, Shurong, et al.
Published: (2024)
by: Yang, Shurong, et al.
Published: (2024)
Dual Teacher Knowledge Distillation with Domain Alignment for Face Anti-spoofing
by: Kong, Zhe, et al.
Published: (2024)
by: Kong, Zhe, et al.
Published: (2024)
Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary Head
by: Yang, Penghui, et al.
Published: (2024)
by: Yang, Penghui, et al.
Published: (2024)
KD-DETR: Knowledge Distillation for Detection Transformer with Consistent Distillation Points Sampling
by: Wang, Yu, et al.
Published: (2022)
by: Wang, Yu, et al.
Published: (2022)
Contrast-Phys: Unsupervised Video-based Remote Physiological Measurement via Spatiotemporal Contrast
by: Sun, Zhaodong, et al.
Published: (2022)
by: Sun, Zhaodong, et al.
Published: (2022)
Contrast-Phys+: Unsupervised and Weakly-supervised Video-based Remote Physiological Measurement via Spatiotemporal Contrast
by: Sun, Zhaodong, et al.
Published: (2023)
by: Sun, Zhaodong, et al.
Published: (2023)
UnMA-CapSumT: Unified and Multi-Head Attention-driven Caption Summarization Transformer
by: Sharma, Dhruv, et al.
Published: (2024)
by: Sharma, Dhruv, et al.
Published: (2024)
Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation
by: Cui, Xiao, et al.
Published: (2025)
by: Cui, Xiao, et al.
Published: (2025)
Soft Knowledge Distillation with Multi-Dimensional Cross-Net Attention for Image Restoration Models Compression
by: Zhang, Yongheng, et al.
Published: (2025)
by: Zhang, Yongheng, et al.
Published: (2025)
ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge Distillation
by: Lan, Qizhen, et al.
Published: (2025)
by: Lan, Qizhen, et al.
Published: (2025)
Interactive Multi-Head Self-Attention with Linear Complexity
by: Kang, Hankyul, et al.
Published: (2024)
by: Kang, Hankyul, et al.
Published: (2024)
MoH: Multi-Head Attention as Mixture-of-Head Attention
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
A Transformer-in-Transformer Network Utilizing Knowledge Distillation for Image Recognition
by: Rahman, Dewan Tauhid, et al.
Published: (2025)
by: Rahman, Dewan Tauhid, et al.
Published: (2025)
HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation
by: Zhang, Sha, et al.
Published: (2024)
by: Zhang, Sha, et al.
Published: (2024)
Knowledge Distillation in Vision Transformers: A Critical Review
by: Habib, Gousia, et al.
Published: (2023)
by: Habib, Gousia, et al.
Published: (2023)
Knowledge Distillation via Query Selection for Detection Transformer
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Transformer-Based Dual-Optical Attention Fusion Crowd Head Point Counting and Localization Network
by: Zhou, Fei, et al.
Published: (2025)
by: Zhou, Fei, et al.
Published: (2025)
Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers
by: Liu, Yuhe, et al.
Published: (2026)
by: Liu, Yuhe, et al.
Published: (2026)
Exploring Inconsistent Knowledge Distillation for Object Detection with Data Augmentation
by: Liang, Jiawei, et al.
Published: (2022)
by: Liang, Jiawei, et al.
Published: (2022)
A Progressive Framework of Vision-language Knowledge Distillation and Alignment for Multilingual Scene
by: Zhang, Wenbo, et al.
Published: (2024)
by: Zhang, Wenbo, et al.
Published: (2024)
Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition
by: Garg, Mallika, et al.
Published: (2025)
by: Garg, Mallika, et al.
Published: (2025)
Cross-modulated Attention Transformer for RGBT Tracking
by: Xiao, Yun, et al.
Published: (2024)
by: Xiao, Yun, et al.
Published: (2024)
HTR-JAND: Handwritten Text Recognition with Joint Attention Network and Knowledge Distillation
by: Hamdan, Mohammed, et al.
Published: (2024)
by: Hamdan, Mohammed, et al.
Published: (2024)
Channel Attention-Guided Cross-Modal Knowledge Distillation for Referring Image Segmentation
by: Yang, Chen
Published: (2026)
by: Yang, Chen
Published: (2026)
Similar Items
-
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
by: Zhang, Shen, et al.
Published: (2025) -
MAT: Multi-Range Attention Transformer for Efficient Image Super-Resolution
by: Xie, Chengxing, et al.
Published: (2024) -
Knowledge Distillation via the Target-aware Transformer
by: Lin, Sihao, et al.
Published: (2022) -
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024) -
MegActor-$Σ$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer
by: Yang, Shurong, et al.
Published: (2024)