Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lanzillotta, Giulia, Sarnthein, Felix, Kur, Gil, Hofmann, Thomas, He, Bobby |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Heads collapse, features stay: Why Replay needs big buffers
von: Lanzillotta, Giulia, et al.
Veröffentlicht: (2025)
von: Lanzillotta, Giulia, et al.
Veröffentlicht: (2025)
The Importance of Being Lazy: Scaling Limits of Continual Learning
von: Graldi, Jacopo, et al.
Veröffentlicht: (2025)
von: Graldi, Jacopo, et al.
Veröffentlicht: (2025)
Local vs Global continual learning
von: Lanzillotta, Giulia, et al.
Veröffentlicht: (2024)
von: Lanzillotta, Giulia, et al.
Veröffentlicht: (2024)
Reactivation: Empirical NTK Dynamics Under Task Shifts
von: Liu, Yuzhi, et al.
Veröffentlicht: (2025)
von: Liu, Yuzhi, et al.
Veröffentlicht: (2025)
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
von: Joudaki, Amir, et al.
Veröffentlicht: (2025)
von: Joudaki, Amir, et al.
Veröffentlicht: (2025)
Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
von: Hübotter, Jonas, et al.
Veröffentlicht: (2025)
von: Hübotter, Jonas, et al.
Veröffentlicht: (2025)
Bridging the Gap: Unpacking the Hidden Challenges in Knowledge Distillation for Online Ranking Systems
von: Khani, Nikhil, et al.
Veröffentlicht: (2024)
von: Khani, Nikhil, et al.
Veröffentlicht: (2024)
A Continual and Incremental Learning Approach for TinyML On-device Training Using Dataset Distillation and Model Size Adaption
von: Rüb, Marcus, et al.
Veröffentlicht: (2024)
von: Rüb, Marcus, et al.
Veröffentlicht: (2024)
Prioritize Alignment in Dataset Distillation
von: Li, Zekai, et al.
Veröffentlicht: (2024)
von: Li, Zekai, et al.
Veröffentlicht: (2024)
scDD: Latent Codes Based scRNA-seq Dataset Distillation with Foundation Model Knowledge
von: Yu, Zhen, et al.
Veröffentlicht: (2025)
von: Yu, Zhen, et al.
Veröffentlicht: (2025)
The Role of Teacher Calibration in Knowledge Distillation
von: Kim, Suyoung, et al.
Veröffentlicht: (2025)
von: Kim, Suyoung, et al.
Veröffentlicht: (2025)
Dataset Distillation for Offline Reinforcement Learning
von: Light, Jonathan, et al.
Veröffentlicht: (2024)
von: Light, Jonathan, et al.
Veröffentlicht: (2024)
Dataset Distillation-based Hybrid Federated Learning on Non-IID Data
von: Shi, Xiufang, et al.
Veröffentlicht: (2024)
von: Shi, Xiufang, et al.
Veröffentlicht: (2024)
On the Role of Hidden States of Modern Hopfield Network in Transformer
von: Masumura, Tsubasa, et al.
Veröffentlicht: (2025)
von: Masumura, Tsubasa, et al.
Veröffentlicht: (2025)
Large Language Model Guided Knowledge Distillation for Time Series Anomaly Detection
von: Liu, Chen, et al.
Veröffentlicht: (2024)
von: Liu, Chen, et al.
Veröffentlicht: (2024)
Dynamic Temperature Scheduler for Knowledge Distillation
von: Islam, Sibgat Ul, et al.
Veröffentlicht: (2025)
von: Islam, Sibgat Ul, et al.
Veröffentlicht: (2025)
Membership and Memorization in LLM Knowledge Distillation
von: Zhang, Ziqi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziqi, et al.
Veröffentlicht: (2025)
Graph Knowledge Distillation to Mixture of Experts
von: Rumiantsev, Pavel, et al.
Veröffentlicht: (2024)
von: Rumiantsev, Pavel, et al.
Veröffentlicht: (2024)
Towards Mitigating Architecture Overfitting on Distilled Datasets
von: Zhong, Xuyang, et al.
Veröffentlicht: (2023)
von: Zhong, Xuyang, et al.
Veröffentlicht: (2023)
Grounding and Enhancing Informativeness and Utility in Dataset Distillation
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
Path-Guided Flow Matching for Dataset Distillation
von: Li, Xuhui, et al.
Veröffentlicht: (2026)
von: Li, Xuhui, et al.
Veröffentlicht: (2026)
Memory Self-Regeneration: Uncovering Hidden Knowledge in Unlearned Models
von: Polowczyk, Agnieszka, et al.
Veröffentlicht: (2025)
von: Polowczyk, Agnieszka, et al.
Veröffentlicht: (2025)
Low-Dimensional Federated Knowledge Graph Embedding via Knowledge Distillation
von: Zhang, Xiaoxiong, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoxiong, et al.
Veröffentlicht: (2024)
Role of Mixup in Topological Persistence Based Knowledge Distillation for Wearable Sensor Data
von: Jeon, Eun Som, et al.
Veröffentlicht: (2025)
von: Jeon, Eun Som, et al.
Veröffentlicht: (2025)
Knowledge Distillation Must Account for What It Loses
von: Wang, Wenshuo
Veröffentlicht: (2026)
von: Wang, Wenshuo
Veröffentlicht: (2026)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2026)
Collaborative Adaptive Curriculum for Progressive Knowledge Distillation
von: Liu, Jing, et al.
Veröffentlicht: (2026)
von: Liu, Jing, et al.
Veröffentlicht: (2026)
REBEL: Hidden Knowledge Recovery via Evolutionary-Based Evaluation Loop
von: Rybak, Patryk, et al.
Veröffentlicht: (2026)
von: Rybak, Patryk, et al.
Veröffentlicht: (2026)
Revisiting Catastrophic Forgetting in Continual Knowledge Graph Embedding
von: Pons, Gerard, et al.
Veröffentlicht: (2026)
von: Pons, Gerard, et al.
Veröffentlicht: (2026)
Teach Harder, Learn Poorer: Rethinking Hard Sample Distillation for GNN-to-MLP Knowledge Distillation
von: Wu, Lirong, et al.
Veröffentlicht: (2024)
von: Wu, Lirong, et al.
Veröffentlicht: (2024)
A Functional Perspective on Knowledge Distillation in Neural Networks
von: Mason-Williams, Israel, et al.
Veröffentlicht: (2025)
von: Mason-Williams, Israel, et al.
Veröffentlicht: (2025)
Model Merging via Multi-Teacher Knowledge Distillation
von: Dalili, Seyed Arshan, et al.
Veröffentlicht: (2025)
von: Dalili, Seyed Arshan, et al.
Veröffentlicht: (2025)
Balancing Knowledge Distillation for Imbalance Learning with Bilevel Optimization
von: Nguyen, Anh B. H., et al.
Veröffentlicht: (2026)
von: Nguyen, Anh B. H., et al.
Veröffentlicht: (2026)
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
von: Bick, Aviv, et al.
Veröffentlicht: (2024)
von: Bick, Aviv, et al.
Veröffentlicht: (2024)
Online Adversarial Knowledge Distillation for Graph Neural Networks
von: Wang, Can, et al.
Veröffentlicht: (2021)
von: Wang, Can, et al.
Veröffentlicht: (2021)
Practical Insights into Knowledge Distillation for Pre-Trained Models
von: Alballa, Norah, et al.
Veröffentlicht: (2024)
von: Alballa, Norah, et al.
Veröffentlicht: (2024)
Rethinking Momentum Knowledge Distillation in Online Continual Learning
von: Michel, Nicolas, et al.
Veröffentlicht: (2023)
von: Michel, Nicolas, et al.
Veröffentlicht: (2023)
Consistently Informative Soft-Label Temperature for Knowledge Distillation
von: Luong, Hoang-Chau, et al.
Veröffentlicht: (2026)
von: Luong, Hoang-Chau, et al.
Veröffentlicht: (2026)
Cooperative Knowledge Distillation: A Learner Agnostic Approach
von: Livanos, Michael, et al.
Veröffentlicht: (2024)
von: Livanos, Michael, et al.
Veröffentlicht: (2024)
Fair Dataset Distillation via Cross-Group Barycenter Alignment
von: Moslemi, Mohammad Hossein, et al.
Veröffentlicht: (2026)
von: Moslemi, Mohammad Hossein, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Heads collapse, features stay: Why Replay needs big buffers
von: Lanzillotta, Giulia, et al.
Veröffentlicht: (2025) -
The Importance of Being Lazy: Scaling Limits of Continual Learning
von: Graldi, Jacopo, et al.
Veröffentlicht: (2025) -
Local vs Global continual learning
von: Lanzillotta, Giulia, et al.
Veröffentlicht: (2024) -
Reactivation: Empirical NTK Dynamics Under Task Shifts
von: Liu, Yuzhi, et al.
Veröffentlicht: (2025) -
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
von: Joudaki, Amir, et al.
Veröffentlicht: (2025)