Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Ruihan, Shao, Pengpeng, Wen, Zhengqi, Wu, Jinyang, Feng, Mingkuan, Yang, Shuo, Zhang, Chu Yuan, Tao, Jianhua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing
by: Jin, Ruihan, et al.
Published: (2025)
by: Jin, Ruihan, et al.
Published: (2025)
DReSS: Data-driven Regularized Structured Streamlining for Large Language Models
by: Feng, Mingkuan, et al.
Published: (2025)
by: Feng, Mingkuan, et al.
Published: (2025)
Two-Stage Regularization-Based Structured Pruning for LLMs
by: Feng, Mingkuan, et al.
Published: (2025)
by: Feng, Mingkuan, et al.
Published: (2025)
AStar: Boosting Multimodal Reasoning with Automated Structured Thinking
by: Wu, Jinyang, et al.
Published: (2025)
by: Wu, Jinyang, et al.
Published: (2025)
Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models
by: Wu, Jinyang, et al.
Published: (2024)
by: Wu, Jinyang, et al.
Published: (2024)
Atlas: Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning
by: Wu, Jinyang, et al.
Published: (2026)
by: Wu, Jinyang, et al.
Published: (2026)
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
by: Wu, Jinyang, et al.
Published: (2025)
by: Wu, Jinyang, et al.
Published: (2025)
Can large language models understand uncommon meanings of common words?
by: Wu, Jinyang, et al.
Published: (2024)
by: Wu, Jinyang, et al.
Published: (2024)
Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning
by: Wu, Jinyang, et al.
Published: (2026)
by: Wu, Jinyang, et al.
Published: (2026)
Beyond Examples: High-level Automated Reasoning Paradigm in In-Context Learning via MCTS
by: Wu, Jinyang, et al.
Published: (2024)
by: Wu, Jinyang, et al.
Published: (2024)
Fake News Detection and Manipulation Reasoning via Large Vision-Language Models
by: Jin, Ruihan, et al.
Published: (2024)
by: Jin, Ruihan, et al.
Published: (2024)
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles
by: Wu, Jinyang, et al.
Published: (2026)
by: Wu, Jinyang, et al.
Published: (2026)
A Survey on Knowledge Distillation of Large Language Models
by: Xu, Xiaohan, et al.
Published: (2024)
by: Xu, Xiaohan, et al.
Published: (2024)
KS-LLM: Knowledge Selection of Large Language Models with Evidence Document for Question Answering
by: Zheng, Xinxin, et al.
Published: (2024)
by: Zheng, Xinxin, et al.
Published: (2024)
From Deferral to Learning: Online In-Context Knowledge Distillation for LLM Cascades
by: Wu, Yu, et al.
Published: (2025)
by: Wu, Yu, et al.
Published: (2025)
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
by: Yuan, Zhonghang, et al.
Published: (2026)
by: Yuan, Zhonghang, et al.
Published: (2026)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' Memory
by: Tao, Xingjian, et al.
Published: (2024)
by: Tao, Xingjian, et al.
Published: (2024)
Exploring and Enhancing the Transfer of Distribution in Knowledge Distillation for Autoregressive Language Models
by: Rao, Jun, et al.
Published: (2024)
by: Rao, Jun, et al.
Published: (2024)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Progressive Distillation Based on Masked Generation Feature Method for Knowledge Graph Completion
by: Fan, Cunhang, et al.
Published: (2024)
by: Fan, Cunhang, et al.
Published: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
3M-Health: Multimodal Multi-Teacher Knowledge Distillation for Mental Health Detection
by: Cabral, Rina Carines, et al.
Published: (2024)
by: Cabral, Rina Carines, et al.
Published: (2024)
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
by: Panigrahi, Abhishek, et al.
Published: (2025)
by: Panigrahi, Abhishek, et al.
Published: (2025)
SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models
by: Koo, Jahyun, et al.
Published: (2024)
by: Koo, Jahyun, et al.
Published: (2024)
DDK: Distilling Domain Knowledge for Efficient Large Language Models
by: Liu, Jiaheng, et al.
Published: (2024)
by: Liu, Jiaheng, et al.
Published: (2024)
Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge
by: Feng, Yiyang, et al.
Published: (2026)
by: Feng, Yiyang, et al.
Published: (2026)
Warmup-Distill: Bridge the Distribution Mismatch between Teacher and Student before Knowledge Distillation
by: Sun, Zengkui, et al.
Published: (2025)
by: Sun, Zengkui, et al.
Published: (2025)
Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
by: Datta, Joyeeta, et al.
Published: (2025)
by: Datta, Joyeeta, et al.
Published: (2025)
Large Scale Knowledge Washing
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
by: Bijoy, Mehedi Hasan, et al.
Published: (2025)
by: Bijoy, Mehedi Hasan, et al.
Published: (2025)
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
by: Fan, Cunhang, et al.
Published: (2023)
by: Fan, Cunhang, et al.
Published: (2023)
Enhancing Knowledge Distillation for LLMs with Response-Priming Prompting
by: Goyal, Vijay, et al.
Published: (2024)
by: Goyal, Vijay, et al.
Published: (2024)
LoCa: Logit Calibration for Knowledge Distillation
by: Yang, Runming, et al.
Published: (2024)
by: Yang, Runming, et al.
Published: (2024)
Distilling Rule-based Knowledge into Large Language Models
by: Yang, Wenkai, et al.
Published: (2023)
by: Yang, Wenkai, et al.
Published: (2023)
SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Adapter-based Selective Knowledge Distillation for Federated Multi-domain Meeting Summarization
by: Feng, Xiachong, et al.
Published: (2023)
by: Feng, Xiachong, et al.
Published: (2023)
LLM-Guided Knowledge Distillation for Temporal Knowledge Graph Reasoning
by: Xing, Wang, et al.
Published: (2026)
by: Xing, Wang, et al.
Published: (2026)
Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery
by: Yang, Chaoqun, et al.
Published: (2026)
by: Yang, Chaoqun, et al.
Published: (2026)
Efficient Knowledge Injection in LLMs via Self-Distillation
by: Kujanpää, Kalle, et al.
Published: (2024)
by: Kujanpää, Kalle, et al.
Published: (2024)
Similar Items
-
RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing
by: Jin, Ruihan, et al.
Published: (2025) -
DReSS: Data-driven Regularized Structured Streamlining for Large Language Models
by: Feng, Mingkuan, et al.
Published: (2025) -
Two-Stage Regularization-Based Structured Pruning for LLMs
by: Feng, Mingkuan, et al.
Published: (2025) -
AStar: Boosting Multimodal Reasoning with Automated Structured Thinking
by: Wu, Jinyang, et al.
Published: (2025) -
Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models
by: Wu, Jinyang, et al.
Published: (2024)