Put Teacher in Student's Shoes: Cross-Distillation for Ultra-compact Model Compression Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Maolin, Chu, Jun, Xie, Sicong, Zang, Xiaoling, Zhao, Yao, Zhong, Wenliang, Zhao, Xiangyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Let's Put Ourselves in Sally's Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models
von: Shinoda, Kazutoshi, et al.
Veröffentlicht: (2025)
von: Shinoda, Kazutoshi, et al.
Veröffentlicht: (2025)
Putting People in LLMs' Shoes: Generating Better Answers via Question Rewriter
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation
von: Xia, Mingxuan, et al.
Veröffentlicht: (2025)
von: Xia, Mingxuan, et al.
Veröffentlicht: (2025)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
Large Multimodal Model Compression via Efficient Pruning and Distillation at AntGroup
von: Wang, Maolin, et al.
Veröffentlicht: (2023)
von: Wang, Maolin, et al.
Veröffentlicht: (2023)
MTA: A Merge-then-Adapt Framework for Personalized Large Language Model
von: Li, Xiaopeng, et al.
Veröffentlicht: (2025)
von: Li, Xiaopeng, et al.
Veröffentlicht: (2025)
Chapter Putting Students in Their Place
von: Cox, Troy | https://orcid.org/0000-0001-9379-5102, et al.
Veröffentlicht: (2026)
von: Cox, Troy | https://orcid.org/0000-0001-9379-5102, et al.
Veröffentlicht: (2026)
Chapter Putting Students in Their Place
von: Cox, Troy | https://orcid.org/0000-0001-9379-5102, et al.
Veröffentlicht: (2026)
von: Cox, Troy | https://orcid.org/0000-0001-9379-5102, et al.
Veröffentlicht: (2026)
Bridging Relevance and Reasoning: Rationale Distillation in Retrieval-Augmented Generation
von: Jia, Pengyue, et al.
Veröffentlicht: (2024)
von: Jia, Pengyue, et al.
Veröffentlicht: (2024)
Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
von: Zang, Jianxiang
Veröffentlicht: (2025)
von: Zang, Jianxiang
Veröffentlicht: (2025)
Warmup-Distill: Bridge the Distribution Mismatch between Teacher and Student before Knowledge Distillation
von: Sun, Zengkui, et al.
Veröffentlicht: (2025)
von: Sun, Zengkui, et al.
Veröffentlicht: (2025)
Aligning Teacher with Student Preferences for Tailored Training Data Generation
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
Job Skill Extraction via LLM-Centric Multi-Module Framework
von: Li, Guojing, et al.
Veröffentlicht: (2026)
von: Li, Guojing, et al.
Veröffentlicht: (2026)
Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
Neural-Symbolic Collaborative Distillation: Advancing Small Language Models for Complex Reasoning Tasks
von: Liao, Huanxuan, et al.
Veröffentlicht: (2024)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2024)
Large Language Models are In-context Teachers for Knowledge Reasoning
von: Zhao, Jiachen, et al.
Veröffentlicht: (2023)
von: Zhao, Jiachen, et al.
Veröffentlicht: (2023)
MetaLoRA: Tensor-Enhanced Adaptive Low-Rank Fine-tuning
von: Wang, Maolin, et al.
Veröffentlicht: (2025)
von: Wang, Maolin, et al.
Veröffentlicht: (2025)
Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
Prompt Recursive Search: A Living Framework with Adaptive Growth in LLM Auto-Prompting
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs
von: Jin, Ruihan, et al.
Veröffentlicht: (2026)
von: Jin, Ruihan, et al.
Veröffentlicht: (2026)
YODA: Teacher-Student Progressive Learning for Language Models
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
Weight-Inherited Distillation for Task-Agnostic BERT Compression
von: Wu, Taiqiang, et al.
Veröffentlicht: (2023)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2023)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
From Images to Words: Efficient Cross-Modal Knowledge Distillation to Language Models from Black-box Teachers
von: Sengupta, Ayan, et al.
Veröffentlicht: (2026)
von: Sengupta, Ayan, et al.
Veröffentlicht: (2026)
Who Taught You That? Tracing Teachers in Model Distillation
von: Wadhwa, Somin, et al.
Veröffentlicht: (2025)
von: Wadhwa, Somin, et al.
Veröffentlicht: (2025)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
von: Zhao, Siyan, et al.
Veröffentlicht: (2026)
von: Zhao, Siyan, et al.
Veröffentlicht: (2026)
Learning User Interests via Reasoning and Distillation for Cross-Domain News Recommendation
von: Zhu, Mengdan, et al.
Veröffentlicht: (2026)
von: Zhu, Mengdan, et al.
Veröffentlicht: (2026)
"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework
von: Cui, Jin, et al.
Veröffentlicht: (2026)
von: Cui, Jin, et al.
Veröffentlicht: (2026)
MiniDisc: Minimal Distillation Schedule for Language Model Compression
von: Zhang, Chen, et al.
Veröffentlicht: (2022)
von: Zhang, Chen, et al.
Veröffentlicht: (2022)
On Teacher Hacking in Language Model Distillation
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
Feature Alignment-Based Knowledge Distillation for Efficient Compression of Large Language Models
von: Wang, Shuo, et al.
Veröffentlicht: (2024)
von: Wang, Shuo, et al.
Veröffentlicht: (2024)
DiaCDM: Cognitive Diagnosis in Teacher-Student Dialogues using the Initiation-Response-Evaluation Framework
von: Jia, Rui, et al.
Veröffentlicht: (2025)
von: Jia, Rui, et al.
Veröffentlicht: (2025)
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
SEARCH-R: Structured Entity-Aware Retrieval with Chain-of-Reasoning Navigator for Multi-hop Question Answering
von: Fu, Yuqing, et al.
Veröffentlicht: (2026)
von: Fu, Yuqing, et al.
Veröffentlicht: (2026)
When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
von: Tuchinda, Pume, et al.
Veröffentlicht: (2025)
von: Tuchinda, Pume, et al.
Veröffentlicht: (2025)
SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
von: Li, Zichong, et al.
Veröffentlicht: (2025)
von: Li, Zichong, et al.
Veröffentlicht: (2025)
Behavior-Equivalent Token: Single-Token Replacement for Long Prompts in LLMs
von: Dong, Jiancheng, et al.
Veröffentlicht: (2025)
von: Dong, Jiancheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Let's Put Ourselves in Sally's Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models
von: Shinoda, Kazutoshi, et al.
Veröffentlicht: (2025) -
Putting People in LLMs' Shoes: Generating Better Answers via Question Rewriter
von: Chen, Junhao, et al.
Veröffentlicht: (2024) -
Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation
von: Xia, Mingxuan, et al.
Veröffentlicht: (2025) -
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
von: Yang, Haoqi, et al.
Veröffentlicht: (2025) -
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)