Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers
Fuente:
arXiv
Guardado en:
| Autores principales: | Bing, Zhaodong, Li, Linze, Liang, Jiajun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
por: Zhang, Shen, et al.
Publicado: (2025)
por: Zhang, Shen, et al.
Publicado: (2025)
MAT: Multi-Range Attention Transformer for Efficient Image Super-Resolution
por: Xie, Chengxing, et al.
Publicado: (2024)
por: Xie, Chengxing, et al.
Publicado: (2024)
Knowledge Distillation via the Target-aware Transformer
por: Lin, Sihao, et al.
Publicado: (2022)
por: Lin, Sihao, et al.
Publicado: (2022)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
por: Zhang, Tianxiao, et al.
Publicado: (2024)
por: Zhang, Tianxiao, et al.
Publicado: (2024)
MegActor-$Σ$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer
por: Yang, Shurong, et al.
Publicado: (2024)
por: Yang, Shurong, et al.
Publicado: (2024)
Frequency Attention for Knowledge Distillation
por: Pham, Cuong, et al.
Publicado: (2024)
por: Pham, Cuong, et al.
Publicado: (2024)
Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention into Convolutions
por: Dong, ZiYi, et al.
Publicado: (2025)
por: Dong, ZiYi, et al.
Publicado: (2025)
CrossKD: Cross-Head Knowledge Distillation for Object Detection
por: Wang, Jiabao, et al.
Publicado: (2023)
por: Wang, Jiabao, et al.
Publicado: (2023)
MoKD: Multi-Task Optimization for Knowledge Distillation
por: Hayder, Zeeshan, et al.
Publicado: (2025)
por: Hayder, Zeeshan, et al.
Publicado: (2025)
Asymmetric Decision-Making in Online Knowledge Distillation:Unifying Consensus and Divergence
por: Chen, Zhaowei, et al.
Publicado: (2025)
por: Chen, Zhaowei, et al.
Publicado: (2025)
The Velocity Deficit: Initial Energy Injection for Flow Matching
por: Li, Linze, et al.
Publicado: (2026)
por: Li, Linze, et al.
Publicado: (2026)
Efficient Temporal Sentence Grounding in Videos with Multi-Teacher Knowledge Distillation
por: Liang, Renjie, et al.
Publicado: (2023)
por: Liang, Renjie, et al.
Publicado: (2023)
Rethinking Centered Kernel Alignment in Knowledge Distillation
por: Zhou, Zikai, et al.
Publicado: (2024)
por: Zhou, Zikai, et al.
Publicado: (2024)
GMAR: Gradient-Driven Multi-Head Attention Rollout for Vision Transformer Interpretability
por: Jo, Sehyeong, et al.
Publicado: (2025)
por: Jo, Sehyeong, et al.
Publicado: (2025)
MHAFF: Multi-Head Attention Feature Fusion of CNN and Transformer for Cattle Identification
por: Dulal, Rabin, et al.
Publicado: (2025)
por: Dulal, Rabin, et al.
Publicado: (2025)
EVOKE: Emotion Enabled Virtual Avatar Mapping Using Optimized Knowledge Distillation
por: Nadeem, Maryam, et al.
Publicado: (2024)
por: Nadeem, Maryam, et al.
Publicado: (2024)
MegActor: Harness the Power of Raw Video for Vivid Portrait Animation
por: Yang, Shurong, et al.
Publicado: (2024)
por: Yang, Shurong, et al.
Publicado: (2024)
Dual Teacher Knowledge Distillation with Domain Alignment for Face Anti-spoofing
por: Kong, Zhe, et al.
Publicado: (2024)
por: Kong, Zhe, et al.
Publicado: (2024)
Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary Head
por: Yang, Penghui, et al.
Publicado: (2024)
por: Yang, Penghui, et al.
Publicado: (2024)
KD-DETR: Knowledge Distillation for Detection Transformer with Consistent Distillation Points Sampling
por: Wang, Yu, et al.
Publicado: (2022)
por: Wang, Yu, et al.
Publicado: (2022)
Contrast-Phys: Unsupervised Video-based Remote Physiological Measurement via Spatiotemporal Contrast
por: Sun, Zhaodong, et al.
Publicado: (2022)
por: Sun, Zhaodong, et al.
Publicado: (2022)
Contrast-Phys+: Unsupervised and Weakly-supervised Video-based Remote Physiological Measurement via Spatiotemporal Contrast
por: Sun, Zhaodong, et al.
Publicado: (2023)
por: Sun, Zhaodong, et al.
Publicado: (2023)
UnMA-CapSumT: Unified and Multi-Head Attention-driven Caption Summarization Transformer
por: Sharma, Dhruv, et al.
Publicado: (2024)
por: Sharma, Dhruv, et al.
Publicado: (2024)
Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation
por: Cui, Xiao, et al.
Publicado: (2025)
por: Cui, Xiao, et al.
Publicado: (2025)
Soft Knowledge Distillation with Multi-Dimensional Cross-Net Attention for Image Restoration Models Compression
por: Zhang, Yongheng, et al.
Publicado: (2025)
por: Zhang, Yongheng, et al.
Publicado: (2025)
ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge Distillation
por: Lan, Qizhen, et al.
Publicado: (2025)
por: Lan, Qizhen, et al.
Publicado: (2025)
Interactive Multi-Head Self-Attention with Linear Complexity
por: Kang, Hankyul, et al.
Publicado: (2024)
por: Kang, Hankyul, et al.
Publicado: (2024)
MoH: Multi-Head Attention as Mixture-of-Head Attention
por: Jin, Peng, et al.
Publicado: (2024)
por: Jin, Peng, et al.
Publicado: (2024)
A Transformer-in-Transformer Network Utilizing Knowledge Distillation for Image Recognition
por: Rahman, Dewan Tauhid, et al.
Publicado: (2025)
por: Rahman, Dewan Tauhid, et al.
Publicado: (2025)
HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation
por: Zhang, Sha, et al.
Publicado: (2024)
por: Zhang, Sha, et al.
Publicado: (2024)
Knowledge Distillation in Vision Transformers: A Critical Review
por: Habib, Gousia, et al.
Publicado: (2023)
por: Habib, Gousia, et al.
Publicado: (2023)
Knowledge Distillation via Query Selection for Detection Transformer
por: Liu, Yi, et al.
Publicado: (2024)
por: Liu, Yi, et al.
Publicado: (2024)
Transformer-Based Dual-Optical Attention Fusion Crowd Head Point Counting and Localization Network
por: Zhou, Fei, et al.
Publicado: (2025)
por: Zhou, Fei, et al.
Publicado: (2025)
Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers
por: Liu, Yuhe, et al.
Publicado: (2026)
por: Liu, Yuhe, et al.
Publicado: (2026)
Exploring Inconsistent Knowledge Distillation for Object Detection with Data Augmentation
por: Liang, Jiawei, et al.
Publicado: (2022)
por: Liang, Jiawei, et al.
Publicado: (2022)
A Progressive Framework of Vision-language Knowledge Distillation and Alignment for Multilingual Scene
por: Zhang, Wenbo, et al.
Publicado: (2024)
por: Zhang, Wenbo, et al.
Publicado: (2024)
Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition
por: Garg, Mallika, et al.
Publicado: (2025)
por: Garg, Mallika, et al.
Publicado: (2025)
Cross-modulated Attention Transformer for RGBT Tracking
por: Xiao, Yun, et al.
Publicado: (2024)
por: Xiao, Yun, et al.
Publicado: (2024)
HTR-JAND: Handwritten Text Recognition with Joint Attention Network and Knowledge Distillation
por: Hamdan, Mohammed, et al.
Publicado: (2024)
por: Hamdan, Mohammed, et al.
Publicado: (2024)
Channel Attention-Guided Cross-Modal Knowledge Distillation for Referring Image Segmentation
por: Yang, Chen
Publicado: (2026)
por: Yang, Chen
Publicado: (2026)
Ejemplares similares
-
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
por: Zhang, Shen, et al.
Publicado: (2025) -
MAT: Multi-Range Attention Transformer for Efficient Image Super-Resolution
por: Xie, Chengxing, et al.
Publicado: (2024) -
Knowledge Distillation via the Target-aware Transformer
por: Lin, Sihao, et al.
Publicado: (2022) -
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
por: Zhang, Tianxiao, et al.
Publicado: (2024) -
MegActor-$Σ$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer
por: Yang, Shurong, et al.
Publicado: (2024)