Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary Head
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Penghui, Zong, Chen-Chen, Huang, Sheng-Jun, Feng, Lei, An, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Label Knowledge Distillation
by: Yang, Penghui, et al.
Published: (2023)
by: Yang, Penghui, et al.
Published: (2023)
Adaptive Temperature Based on Logits Correlation in Knowledge Distillation
by: Matsuyama, Kazuhiro, et al.
Published: (2025)
by: Matsuyama, Kazuhiro, et al.
Published: (2025)
Decoupling Dark Knowledge via Block-wise Logit Distillation for Feature-level Alignment
by: Yu, Chengting, et al.
Published: (2024)
by: Yu, Chengting, et al.
Published: (2024)
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
by: Kang, Seongjae, et al.
Published: (2025)
by: Kang, Seongjae, et al.
Published: (2025)
Aligning Logits Generatively for Principled Black-Box Knowledge Distillation
by: Ma, Jing, et al.
Published: (2022)
by: Ma, Jing, et al.
Published: (2022)
MoH: Multi-Head Attention as Mixture-of-Head Attention
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Logit-Based Losses Limit the Effectiveness of Feature Knowledge Distillation
by: Cooper, Nicholas, et al.
Published: (2025)
by: Cooper, Nicholas, et al.
Published: (2025)
Knowledge Distillation with Refined Logits
by: Sun, Wujie, et al.
Published: (2024)
by: Sun, Wujie, et al.
Published: (2024)
LokiTalk: Learning Fine-Grained and Generalizable Correspondences to Enhance NeRF-based Talking Head Synthesis
by: Li, Tianqi, et al.
Published: (2024)
by: Li, Tianqi, et al.
Published: (2024)
Semi-Supervised Learning with Multi-Head Co-Training
by: Chen, Mingcai, et al.
Published: (2021)
by: Chen, Mingcai, et al.
Published: (2021)
Self-supervised Fusarium Head Blight Detection with Hyperspectral Image and Feature Mining
by: Lin, Yu-Fan, et al.
Published: (2024)
by: Lin, Yu-Fan, et al.
Published: (2024)
Revisiting Unknowns: Towards Effective and Efficient Open-Set Active Learning
by: Zong, Chen-Chen, et al.
Published: (2026)
by: Zong, Chen-Chen, et al.
Published: (2026)
Dual-granularity Sinkhorn Distillation for Enhanced Learning from Long-tailed Noisy Data
by: Hong, Feng, et al.
Published: (2025)
by: Hong, Feng, et al.
Published: (2025)
FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation
by: Ma, Yuting, et al.
Published: (2024)
by: Ma, Yuting, et al.
Published: (2024)
Local Dense Logit Relations for Enhanced Knowledge Distillation
by: Xu, Liuchi, et al.
Published: (2025)
by: Xu, Liuchi, et al.
Published: (2025)
Distilling Knowledge from Heterogeneous Architectures for Semantic Segmentation
by: Huang, Yanglin, et al.
Published: (2025)
by: Huang, Yanglin, et al.
Published: (2025)
MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
Generative Head-Mounted Camera Captures for Photorealistic Avatars
by: Bai, Shaojie, et al.
Published: (2025)
by: Bai, Shaojie, et al.
Published: (2025)
SimCast: Enhancing Precipitation Nowcasting with Short-to-Long Term Knowledge Distillation
by: Yin, Yifang, et al.
Published: (2025)
by: Yin, Yifang, et al.
Published: (2025)
HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
by: Xu, Yu, et al.
Published: (2024)
by: Xu, Yu, et al.
Published: (2024)
Investigating the Benefits of Projection Head for Representation Learning
by: Xue, Yihao, et al.
Published: (2024)
by: Xue, Yihao, et al.
Published: (2024)
Uncertainty-Aware Dual-Student Knowledge Distillation for Efficient Image Classification
by: Gore, Aakash, et al.
Published: (2025)
by: Gore, Aakash, et al.
Published: (2025)
Leaner Transformers: More Heads, Less Depth
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
EDocNet: Efficient Datasheet Layout Analysis Based on Focus and Global Knowledge Distillation
by: Chen, Hong Cai, et al.
Published: (2025)
by: Chen, Hong Cai, et al.
Published: (2025)
Attention-ResUNet for Automated Fetal Head Segmentation
by: Bhilwarawala, Ammar, et al.
Published: (2026)
by: Bhilwarawala, Ammar, et al.
Published: (2026)
LAF-YOLOv10 with Partial Convolution Backbone, Attention-Guided Feature Pyramid, Auxiliary P2 Head, and Wise-IoU Loss for Small Object Detection in Drone Aerial Imagery
by: Farooqui, Sohail Ali, et al.
Published: (2026)
by: Farooqui, Sohail Ali, et al.
Published: (2026)
Conditional Pseudo-Supervised Contrast for Data-Free Knowledge Distillation
by: Shao, Renrong, et al.
Published: (2025)
by: Shao, Renrong, et al.
Published: (2025)
CrossKD: Cross-Head Knowledge Distillation for Object Detection
by: Wang, Jiabao, et al.
Published: (2023)
by: Wang, Jiabao, et al.
Published: (2023)
Knowledge Distillation Based on Transformed Teacher Matching
by: Zheng, Kaixiang, et al.
Published: (2024)
by: Zheng, Kaixiang, et al.
Published: (2024)
Tilt your Head: Activating the Hidden Spatial-Invariance of Classifiers
by: Schmidt, Johann, et al.
Published: (2024)
by: Schmidt, Johann, et al.
Published: (2024)
Mathematical Foundation and Corrections for Full Range Head Pose Estimation
by: Hu, Huei-Chung, et al.
Published: (2024)
by: Hu, Huei-Chung, et al.
Published: (2024)
SUPER: Selfie Undistortion and Head Pose Editing with Identity Preservation
by: Karpikova, Polina, et al.
Published: (2024)
by: Karpikova, Polina, et al.
Published: (2024)
Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
by: Gabetni, Firas, et al.
Published: (2025)
by: Gabetni, Firas, et al.
Published: (2025)
EOL: Transductive Few-Shot Open-Set Recognition by Enhancing Outlier Logits
by: Ochal, Mateusz, et al.
Published: (2024)
by: Ochal, Mateusz, et al.
Published: (2024)
Logit Standardization in Knowledge Distillation
by: Sun, Shangquan, et al.
Published: (2024)
by: Sun, Shangquan, et al.
Published: (2024)
m2mKD: Module-to-Module Knowledge Distillation for Modular Transformers
by: Lo, Ka Man, et al.
Published: (2024)
by: Lo, Ka Man, et al.
Published: (2024)
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
by: Lin, Feng, et al.
Published: (2025)
by: Lin, Feng, et al.
Published: (2025)
Highlight Every Step: Knowledge Distillation via Collaborative Teaching
by: Zhao, Haoran, et al.
Published: (2019)
by: Zhao, Haoran, et al.
Published: (2019)
SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation
by: Wu, Zhuguanyu, et al.
Published: (2026)
by: Wu, Zhuguanyu, et al.
Published: (2026)
Identity Preserving 3D Head Stylization with Multiview Score Distillation
by: Bilecen, Bahri Batuhan, et al.
Published: (2024)
by: Bilecen, Bahri Batuhan, et al.
Published: (2024)
Similar Items
-
Multi-Label Knowledge Distillation
by: Yang, Penghui, et al.
Published: (2023) -
Adaptive Temperature Based on Logits Correlation in Knowledge Distillation
by: Matsuyama, Kazuhiro, et al.
Published: (2025) -
Decoupling Dark Knowledge via Block-wise Logit Distillation for Feature-level Alignment
by: Yu, Chengting, et al.
Published: (2024) -
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
by: Kang, Seongjae, et al.
Published: (2025) -
Aligning Logits Generatively for Principled Black-Box Knowledge Distillation
by: Ma, Jing, et al.
Published: (2022)