Knowledge Distillation Must Account for What It Loses
Fuente:
arXiv
Saved in:
| Main Author: | Wang, Wenshuo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trustworthy AI Must Account for Interactions
by: Cresswell, Jesse C.
Published: (2025)
by: Cresswell, Jesse C.
Published: (2025)
S^2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting
by: Wang, Wenshuo, et al.
Published: (2025)
by: Wang, Wenshuo, et al.
Published: (2025)
Simulating the Unseen: Crash Prediction Must Learn from What Did Not Happen
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
Gradient-Informed Temporal Sampling Improves Rollout Accuracy in PDE Surrogate Training
by: Wang, Wenshuo, et al.
Published: (2026)
by: Wang, Wenshuo, et al.
Published: (2026)
Derived Fields Preserve Fine-Scale Detail in Budgeted Neural Simulators
by: Wang, Wenshuo, et al.
Published: (2026)
by: Wang, Wenshuo, et al.
Published: (2026)
Posterior-First Neural PDE Simulation: Inferring Hidden Problem State from a Single Field
by: Wang, Wenshuo, et al.
Published: (2026)
by: Wang, Wenshuo, et al.
Published: (2026)
Losing is for Cherishing: Data Valuation Based on Machine Unlearning and Shapley Value
by: Ma, Le, et al.
Published: (2025)
by: Ma, Le, et al.
Published: (2025)
Large Language Models Must Be Taught to Know What They Don't Know
by: Kapoor, Sanyam, et al.
Published: (2024)
by: Kapoor, Sanyam, et al.
Published: (2024)
Towards Resource-Efficient LLMs: End-to-End Energy Accounting of Distillation Pipelines
by: Lambert, Katherine, et al.
Published: (2026)
by: Lambert, Katherine, et al.
Published: (2026)
Fuse It or Lose It: Deep Fusion for Multimodal Simulation-Based Inference
by: Schmitt, Marvin, et al.
Published: (2023)
by: Schmitt, Marvin, et al.
Published: (2023)
Online Adversarial Knowledge Distillation for Graph Neural Networks
by: Wang, Can, et al.
Published: (2021)
by: Wang, Can, et al.
Published: (2021)
Do Neural Networks Lose Plasticity in a Gradually Changing World?
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
Rethinking Momentum Knowledge Distillation in Online Continual Learning
by: Michel, Nicolas, et al.
Published: (2023)
by: Michel, Nicolas, et al.
Published: (2023)
Graph Knowledge Distillation to Mixture of Experts
by: Rumiantsev, Pavel, et al.
Published: (2024)
by: Rumiantsev, Pavel, et al.
Published: (2024)
Dynamic Temperature Scheduler for Knowledge Distillation
by: Islam, Sibgat Ul, et al.
Published: (2025)
by: Islam, Sibgat Ul, et al.
Published: (2025)
Membership and Memorization in LLM Knowledge Distillation
by: Zhang, Ziqi, et al.
Published: (2025)
by: Zhang, Ziqi, et al.
Published: (2025)
DistilCLIP-EEG: Enhancing Epileptic Seizure Detection Through Multi-modal Learning and Knowledge Distillation
by: Wang, Zexin, et al.
Published: (2025)
by: Wang, Zexin, et al.
Published: (2025)
The Cell Must Go On: Agar.io for Continual Reinforcement Learning
by: Mohamed, Mohamed A., et al.
Published: (2025)
by: Mohamed, Mohamed A., et al.
Published: (2025)
Low-Dimensional Federated Knowledge Graph Embedding via Knowledge Distillation
by: Zhang, Xiaoxiong, et al.
Published: (2024)
by: Zhang, Xiaoxiong, et al.
Published: (2024)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
by: Zhang, Zhaoyang, et al.
Published: (2026)
by: Zhang, Zhaoyang, et al.
Published: (2026)
Collaborative Adaptive Curriculum for Progressive Knowledge Distillation
by: Liu, Jing, et al.
Published: (2026)
by: Liu, Jing, et al.
Published: (2026)
Enhancing Knowledge Graph Completion with GNN Distillation and Probabilistic Interaction Modeling
by: Wang, Lingzhi, et al.
Published: (2025)
by: Wang, Lingzhi, et al.
Published: (2025)
Enhancing Transformer with GNN Structural Knowledge via Distillation: A Novel Approach
by: Duan, Zhihua, et al.
Published: (2025)
by: Duan, Zhihua, et al.
Published: (2025)
Position: A Theory of Deep Learning Must Include Compositional Sparsity
by: Danhofer, David A., et al.
Published: (2025)
by: Danhofer, David A., et al.
Published: (2025)
Split Knowledge Distillation for Large Models in IoT: Architecture, Challenges, and Solutions
by: Li, Zuguang, et al.
Published: (2024)
by: Li, Zuguang, et al.
Published: (2024)
What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty
by: Nayebi, Aran
Published: (2026)
by: Nayebi, Aran
Published: (2026)
Teach Harder, Learn Poorer: Rethinking Hard Sample Distillation for GNN-to-MLP Knowledge Distillation
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
Balancing Knowledge Distillation for Imbalance Learning with Bilevel Optimization
by: Nguyen, Anh B. H., et al.
Published: (2026)
by: Nguyen, Anh B. H., et al.
Published: (2026)
Consistently Informative Soft-Label Temperature for Knowledge Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026)
by: Luong, Hoang-Chau, et al.
Published: (2026)
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
by: Bick, Aviv, et al.
Published: (2024)
by: Bick, Aviv, et al.
Published: (2024)
A Functional Perspective on Knowledge Distillation in Neural Networks
by: Mason-Williams, Israel, et al.
Published: (2025)
by: Mason-Williams, Israel, et al.
Published: (2025)
Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
by: Lanzillotta, Giulia, et al.
Published: (2025)
by: Lanzillotta, Giulia, et al.
Published: (2025)
Model Merging via Multi-Teacher Knowledge Distillation
by: Dalili, Seyed Arshan, et al.
Published: (2025)
by: Dalili, Seyed Arshan, et al.
Published: (2025)
Practical Insights into Knowledge Distillation for Pre-Trained Models
by: Alballa, Norah, et al.
Published: (2024)
by: Alballa, Norah, et al.
Published: (2024)
Cooperative Knowledge Distillation: A Learner Agnostic Approach
by: Livanos, Michael, et al.
Published: (2024)
by: Livanos, Michael, et al.
Published: (2024)
Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs, Shaping Diversity
by: Kruszewski, Germán, et al.
Published: (2025)
by: Kruszewski, Germán, et al.
Published: (2025)
Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPs
by: Wu, Taiqiang, et al.
Published: (2023)
by: Wu, Taiqiang, et al.
Published: (2023)
Efficient and Robust Knowledge Distillation from A Stronger Teacher Based on Correlation Matching
by: Niu, Wenqi, et al.
Published: (2024)
by: Niu, Wenqi, et al.
Published: (2024)
The Override Gap: A Magnitude Account of Knowledge Conflict Failure in Hypernetwork-Based Instant LLM Adaptation
by: Cheng, Shuaizhi, et al.
Published: (2026)
by: Cheng, Shuaizhi, et al.
Published: (2026)
Generative Diffusion Prior Distillation for Long-Context Knowledge Transfer
by: Udayangani, Nilushika, et al.
Published: (2026)
by: Udayangani, Nilushika, et al.
Published: (2026)
Similar Items
-
Trustworthy AI Must Account for Interactions
by: Cresswell, Jesse C.
Published: (2025) -
S^2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting
by: Wang, Wenshuo, et al.
Published: (2025) -
Simulating the Unseen: Crash Prediction Must Learn from What Did Not Happen
by: Li, Zihao, et al.
Published: (2025) -
Gradient-Informed Temporal Sampling Improves Rollout Accuracy in PDE Surrogate Training
by: Wang, Wenshuo, et al.
Published: (2026) -
Derived Fields Preserve Fine-Scale Detail in Budgeted Neural Simulators
by: Wang, Wenshuo, et al.
Published: (2026)