Two Heads are Better than One: Distilling Large Language Model Features Into Small Models with Feature Decomposition and Mixture
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Tianhao, Xu, Xinxin, Xu, Weichen, Chen, Jue, Ren, Ruilong, Deng, Bowen, Zhao, Xinyu, Cao, Jian, Cao, Xixin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Two Heads Are Better than One: Simulating Large Transformers with Small Ones
by: Yu, Hantao, et al.
Published: (2025)
by: Yu, Hantao, et al.
Published: (2025)
Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
Two Heads are Better than One: Robust Learning Meets Multi-branch Models
by: Zhang, Zongyuan, et al.
Published: (2022)
by: Zhang, Zongyuan, et al.
Published: (2022)
The Knowledge Microscope: Features as Better Analytical Lenses than Neurons
by: Chen, Yuheng, et al.
Published: (2025)
by: Chen, Yuheng, et al.
Published: (2025)
Feature Distillation is the Better Choice for Model-Heterogeneous Federated Learning
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
Two Heads Better than One: Dual Degradation Representation for Blind Super-Resolution
by: Yuan, Hsuan, et al.
Published: (2025)
by: Yuan, Hsuan, et al.
Published: (2025)
Two Heads are Actually Better than One: Towards Better Adversarial Robustness via Transduction and Rejection
by: Palumbo, Nils, et al.
Published: (2023)
by: Palumbo, Nils, et al.
Published: (2023)
One LMC Is Better than Two
Published: (1973)
Published: (1973)
PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning
by: Zhu, Xuekai, et al.
Published: (2023)
by: Zhu, Xuekai, et al.
Published: (2023)
Two Heads are Better than One: Geometric-Latent Attention for Point Cloud Classification and Segmentation
by: Cuevas-Velasquez, Hanz, et al.
Published: (2021)
by: Cuevas-Velasquez, Hanz, et al.
Published: (2021)
Federated Distillation for Whole Slide Image via Gaussian-Mixture Feature Alignment and Curriculum Integration
by: Jing, Luru, et al.
Published: (2026)
by: Jing, Luru, et al.
Published: (2026)
Two is Better than One: Efficient Ensemble Defense for Robust and Compact Models
by: Jung, Yoojin, et al.
Published: (2025)
by: Jung, Yoojin, et al.
Published: (2025)
Two Heads are Better than One: Nested PoE for Robust Defense Against Multi-Backdoors
by: Graf, Victoria, et al.
Published: (2024)
by: Graf, Victoria, et al.
Published: (2024)
Two Heads Are Better than One: Influencing Preservice Classroom Teachers' Understanding and Practice of Classroom-Library Collaboration
by: Moreillon, Judi
Published: (2008)
by: Moreillon, Judi
Published: (2008)
Two Heads Are Better than One: Model-Weight and Latent-Space Analysis for Federated Learning on Non-iid Data against Poisoning Attacks
by: Lyu, Xingyu, et al.
Published: (2025)
by: Lyu, Xingyu, et al.
Published: (2025)
Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time
by: Li, Jiazheng, et al.
Published: (2025)
by: Li, Jiazheng, et al.
Published: (2025)
Hybrid Attention Model Using Feature Decomposition and Knowledge Distillation for Glucose Forecasting
by: Farahmand, Ebrahim, et al.
Published: (2024)
by: Farahmand, Ebrahim, et al.
Published: (2024)
Meta Co-Training: Two Views are Better than One
by: Rothenberger, Jay C., et al.
Published: (2023)
by: Rothenberger, Jay C., et al.
Published: (2023)
Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
by: Deng, Boyi, et al.
Published: (2026)
by: Deng, Boyi, et al.
Published: (2026)
Suppress Content Shift: Better Diffusion Features via Off-the-Shelf Generation Techniques
by: Meng, Benyuan, et al.
Published: (2024)
by: Meng, Benyuan, et al.
Published: (2024)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Three Heads Are Better Than One: Complementary Experts for Long-Tailed Semi-supervised Learning
by: Ma, Chengcheng, et al.
Published: (2023)
by: Ma, Chengcheng, et al.
Published: (2023)
Innovative Tooth Segmentation Using Hierarchical Features and Bidirectional Sequence Modeling
by: Zhao, Xinxin, et al.
Published: (2026)
by: Zhao, Xinxin, et al.
Published: (2026)
Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
by: Han, Runduo, et al.
Published: (2024)
by: Han, Runduo, et al.
Published: (2024)
Botfip-LLM: An Enhanced Multimodal Scientific Computing Framework Leveraging Knowledge Distillation from Large Language Models
by: Chen, Tianhao, et al.
Published: (2024)
by: Chen, Tianhao, et al.
Published: (2024)
Model Selection and Parameter Estimation of One-Dimensional Gaussian Mixture Models
by: Liu, Xinyu, et al.
Published: (2024)
by: Liu, Xinyu, et al.
Published: (2024)
Two Heads are Better Than One: Team Teaching in the Information Age.
by: Jurena, Donna Phin, et al.
Published: (1997)
by: Jurena, Donna Phin, et al.
Published: (1997)
On the Feature Learning in Diffusion Models
by: Han, Andi, et al.
Published: (2024)
by: Han, Andi, et al.
Published: (2024)
A Conformal Approach to Feature-based Newsvendor under Model Misspecification
by: Cao, Junyu
Published: (2024)
by: Cao, Junyu
Published: (2024)
Distilling Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection
by: Zheng, Haowen, et al.
Published: (2024)
by: Zheng, Haowen, et al.
Published: (2024)
Not All Diffusion Model Activations Have Been Evaluated as Discriminative Features
by: Meng, Benyuan, et al.
Published: (2024)
by: Meng, Benyuan, et al.
Published: (2024)
Improving Neural Surface Reconstruction with Feature Priors from Multi-View Image
by: Ren, Xinlin, et al.
Published: (2024)
by: Ren, Xinlin, et al.
Published: (2024)
Dynamic Texture Transfer using PatchMatch and Transformers
by: Pu, Guo, et al.
Published: (2024)
by: Pu, Guo, et al.
Published: (2024)
Progressive Feature Learning for Realistic Cloth-Changing Gait Recognition
by: Ren, Xuqian, et al.
Published: (2022)
by: Ren, Xuqian, et al.
Published: (2022)
CLIP Brings Better Features to Visual Aesthetics Learners
by: Xu, Liwu, et al.
Published: (2023)
by: Xu, Liwu, et al.
Published: (2023)
Causality-inspired Latent Feature Augmentation for Single Domain Generalization
by: Xu, Jian, et al.
Published: (2024)
by: Xu, Jian, et al.
Published: (2024)
Lamer-SSL: Layer-aware Mixture of LoRA Experts for Continual Multilingual Expansion of Self-supervised Models without Forgetting
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
DAF-Net: A Dual-Branch Feature Decomposition Fusion Network with Domain Adaptive for Infrared and Visible Image Fusion
by: Xu, Jian, et al.
Published: (2024)
by: Xu, Jian, et al.
Published: (2024)
SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation
by: Wang, Qianxu, et al.
Published: (2023)
by: Wang, Qianxu, et al.
Published: (2023)
Done Is Better than Perfect: Unlocking Efficient Reasoning by Structured Multi-Turn Decomposition
by: Zeng, Zihao, et al.
Published: (2025)
by: Zeng, Zihao, et al.
Published: (2025)
Similar Items
-
Two Heads Are Better than One: Simulating Large Transformers with Small Ones
by: Yu, Hantao, et al.
Published: (2025) -
Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning
by: Zhang, Yichi, et al.
Published: (2024) -
Two Heads are Better than One: Robust Learning Meets Multi-branch Models
by: Zhang, Zongyuan, et al.
Published: (2022) -
The Knowledge Microscope: Features as Better Analytical Lenses than Neurons
by: Chen, Yuheng, et al.
Published: (2025) -
Feature Distillation is the Better Choice for Model-Heterogeneous Federated Learning
by: Li, Yichen, et al.
Published: (2025)