Practical Insights into Knowledge Distillation for Pre-Trained Models
Fuente:
arXiv
Saved in:
| Main Authors: | Alballa, Norah, Abdelmoniem, Ahmed M., Canini, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Query-based Knowledge Transfer for Heterogeneous Learning Environments
by: Alballa, Norah, et al.
Published: (2025)
by: Alballa, Norah, et al.
Published: (2025)
Flashback: Understanding and Mitigating Forgetting in Federated Learning
by: Aljahdali, Mohammed, et al.
Published: (2024)
by: Aljahdali, Mohammed, et al.
Published: (2024)
DeepFusion: Accelerating MoE Training via Federated Knowledge Distillation from Heterogeneous Edge Devices
by: Li, Songyuan, et al.
Published: (2026)
by: Li, Songyuan, et al.
Published: (2026)
A Meta-learning based Stacked Regression Approach for Customer Lifetime Value Prediction
by: Gadgil, Karan, et al.
Published: (2023)
by: Gadgil, Karan, et al.
Published: (2023)
Stock Market Price Prediction: A Hybrid LSTM and Sequential Self-Attention based Approach
by: Pardeshi, Karan, et al.
Published: (2023)
by: Pardeshi, Karan, et al.
Published: (2023)
On the Surprising Efficacy of Distillation as an Alternative to Pre-Training Small Models
by: Farhat, Sean, et al.
Published: (2024)
by: Farhat, Sean, et al.
Published: (2024)
Federated Knowledge Transfer Fine-tuning Large Server Model with Resource-Constrained IoT Clients
by: Chen, Shaoyuan, et al.
Published: (2024)
by: Chen, Shaoyuan, et al.
Published: (2024)
Pre-Training Graph Contrastive Masked Autoencoders are Strong Distillers for EEG
by: Wei, Xinxu, et al.
Published: (2024)
by: Wei, Xinxu, et al.
Published: (2024)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
by: Xin, Jihao, et al.
Published: (2026)
by: Xin, Jihao, et al.
Published: (2026)
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
by: Dong, Chengyu, et al.
Published: (2022)
by: Dong, Chengyu, et al.
Published: (2022)
Training Plug-n-Play Knowledge Modules with Deep Context Distillation
by: Caccia, Lucas, et al.
Published: (2025)
by: Caccia, Lucas, et al.
Published: (2025)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
Training Language Models on the Knowledge Graph: Insights on Hallucinations and Their Detectability
by: Hron, Jiri, et al.
Published: (2024)
by: Hron, Jiri, et al.
Published: (2024)
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
by: Bick, Aviv, et al.
Published: (2024)
by: Bick, Aviv, et al.
Published: (2024)
Model Merging via Multi-Teacher Knowledge Distillation
by: Dalili, Seyed Arshan, et al.
Published: (2025)
by: Dalili, Seyed Arshan, et al.
Published: (2025)
Panther: Faster and Cheaper Computations with Randomized Numerical Linear Algebra
by: Seddik, Fahd, et al.
Published: (2026)
by: Seddik, Fahd, et al.
Published: (2026)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
by: Kharrat, Salma, et al.
Published: (2024)
by: Kharrat, Salma, et al.
Published: (2024)
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
by: Hu, Renjun, et al.
Published: (2025)
by: Hu, Renjun, et al.
Published: (2025)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
by: Hong, Fenglu, et al.
Published: (2025)
by: Hong, Fenglu, et al.
Published: (2025)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
by: Wang, Keyu, et al.
Published: (2025)
by: Wang, Keyu, et al.
Published: (2025)
Multi-Stage Knowledge-Distilled VGAE and GAT for Robust Controller-Area-Network Intrusion Detection
by: Frenken, Robert, et al.
Published: (2025)
by: Frenken, Robert, et al.
Published: (2025)
A Time Series Multitask Framework Integrating a Large Language Model, Pre-Trained Time Series Model, and Knowledge Graph
by: Hao, Shule, et al.
Published: (2025)
by: Hao, Shule, et al.
Published: (2025)
Condensed Data Expansion Using Model Inversion for Knowledge Distillation
by: Binici, Kuluhan, et al.
Published: (2024)
by: Binici, Kuluhan, et al.
Published: (2024)
A Survey on Time-Series Pre-Trained Models
by: Ma, Qianli, et al.
Published: (2023)
by: Ma, Qianli, et al.
Published: (2023)
Graph Knowledge Distillation to Mixture of Experts
by: Rumiantsev, Pavel, et al.
Published: (2024)
by: Rumiantsev, Pavel, et al.
Published: (2024)
Dynamic Temperature Scheduler for Knowledge Distillation
by: Islam, Sibgat Ul, et al.
Published: (2025)
by: Islam, Sibgat Ul, et al.
Published: (2025)
Membership and Memorization in LLM Knowledge Distillation
by: Zhang, Ziqi, et al.
Published: (2025)
by: Zhang, Ziqi, et al.
Published: (2025)
Split Knowledge Distillation for Large Models in IoT: Architecture, Challenges, and Solutions
by: Li, Zuguang, et al.
Published: (2024)
by: Li, Zuguang, et al.
Published: (2024)
Model Mimic Attack: Knowledge Distillation for Provably Transferable Adversarial Examples
by: Lukyanov, Kirill, et al.
Published: (2024)
by: Lukyanov, Kirill, et al.
Published: (2024)
Enhancing Knowledge Graph Completion with GNN Distillation and Probabilistic Interaction Modeling
by: Wang, Lingzhi, et al.
Published: (2025)
by: Wang, Lingzhi, et al.
Published: (2025)
Topology Only Pre-Training: Towards Generalised Multi-Domain Graph Models
by: Davies, Alex O., et al.
Published: (2023)
by: Davies, Alex O., et al.
Published: (2023)
Ensemble of Pre-Trained Models for Long-Tailed Trajectory Prediction
by: Thuremella, Divya, et al.
Published: (2025)
by: Thuremella, Divya, et al.
Published: (2025)
KD-GAT: Combining Knowledge Distillation and Graph Attention Transformer for a Controller Area Network Intrusion Detection System
by: Frenken, Robert, et al.
Published: (2025)
by: Frenken, Robert, et al.
Published: (2025)
Low-Dimensional Federated Knowledge Graph Embedding via Knowledge Distillation
by: Zhang, Xiaoxiong, et al.
Published: (2024)
by: Zhang, Xiaoxiong, et al.
Published: (2024)
Robust Knowledge Distillation Based on Feature Variance Against Backdoored Teacher Model
by: Chen, Jinyin, et al.
Published: (2024)
by: Chen, Jinyin, et al.
Published: (2024)
Large Language Model Guided Knowledge Distillation for Time Series Anomaly Detection
by: Liu, Chen, et al.
Published: (2024)
by: Liu, Chen, et al.
Published: (2024)
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
by: Park, Seonghyeon, et al.
Published: (2026)
by: Park, Seonghyeon, et al.
Published: (2026)
Knowledge Distillation Must Account for What It Loses
by: Wang, Wenshuo
Published: (2026)
by: Wang, Wenshuo
Published: (2026)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
by: Zhang, Zhaoyang, et al.
Published: (2026)
by: Zhang, Zhaoyang, et al.
Published: (2026)
Collaborative Adaptive Curriculum for Progressive Knowledge Distillation
by: Liu, Jing, et al.
Published: (2026)
by: Liu, Jing, et al.
Published: (2026)
Similar Items
-
Query-based Knowledge Transfer for Heterogeneous Learning Environments
by: Alballa, Norah, et al.
Published: (2025) -
Flashback: Understanding and Mitigating Forgetting in Federated Learning
by: Aljahdali, Mohammed, et al.
Published: (2024) -
DeepFusion: Accelerating MoE Training via Federated Knowledge Distillation from Heterogeneous Edge Devices
by: Li, Songyuan, et al.
Published: (2026) -
A Meta-learning based Stacked Regression Approach for Customer Lifetime Value Prediction
by: Gadgil, Karan, et al.
Published: (2023) -
Stock Market Price Prediction: A Hybrid LSTM and Sequential Self-Attention based Approach
by: Pardeshi, Karan, et al.
Published: (2023)