Understanding the Gains from Repeated Self-Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Pareek, Divyansh, Du, Simon S., Oh, Sewoong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding the Gain from Data Filtering in Multimodal Contrastive Learning
by: Pareek, Divyansh, et al.
Published: (2025)
by: Pareek, Divyansh, et al.
Published: (2025)
PLeaS -- Merging Models with Permutations and Least Squares
by: Nasery, Anshul, et al.
Published: (2024)
by: Nasery, Anshul, et al.
Published: (2024)
Self Distillation via Iterative Constructive Perturbations
by: Dave, Maheak, et al.
Published: (2025)
by: Dave, Maheak, et al.
Published: (2025)
Randomization Techniques to Mitigate the Risk of Copyright Infringement
by: Chen, Wei-Ning, et al.
Published: (2024)
by: Chen, Wei-Ning, et al.
Published: (2024)
Training AI to be Loyal
by: Oh, Sewoong, et al.
Published: (2025)
by: Oh, Sewoong, et al.
Published: (2025)
Sampling from Your Language Model One Byte at a Time
by: Hayase, Jonathan, et al.
Published: (2025)
by: Hayase, Jonathan, et al.
Published: (2025)
Characterizing the Training Dynamics of Private Fine-tuning with Langevin diffusion
by: Ke, Shuqi, et al.
Published: (2024)
by: Ke, Shuqi, et al.
Published: (2024)
Do Vision--Language Models Understand 3D Scenes or Just Catalogue Objects?
by: Maheshwari, Animesh, et al.
Published: (2026)
by: Maheshwari, Animesh, et al.
Published: (2026)
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?
by: Hayase, Jonathan, et al.
Published: (2024)
by: Hayase, Jonathan, et al.
Published: (2024)
DeepPolar: Inventing Nonlinear Large-Kernel Polar Codes via Deep Learning
by: Hebbar, S Ashwin, et al.
Published: (2024)
by: Hebbar, S Ashwin, et al.
Published: (2024)
Are Robust LLM Fingerprints Adversarially Robust?
by: Nasery, Anshul, et al.
Published: (2025)
by: Nasery, Anshul, et al.
Published: (2025)
S4S: Solving for a Diffusion Model Solver
by: Frankel, Eric, et al.
Published: (2025)
by: Frankel, Eric, et al.
Published: (2025)
Improved Communication-Privacy Trade-offs in $L_2$ Mean Estimation under Streaming Differential Privacy
by: Chen, Wei-Ning, et al.
Published: (2024)
by: Chen, Wei-Ning, et al.
Published: (2024)
A Multi-Component AI Framework for Computational Psychology: From Robust Predictive Modeling to Deployed Generative Dialogue
by: Pareek, Anant
Published: (2025)
by: Pareek, Anant
Published: (2025)
DPZero: Private Fine-Tuning of Language Models without Backpropagation
by: Zhang, Liang, et al.
Published: (2023)
by: Zhang, Liang, et al.
Published: (2023)
SuperBPE: Space Travel for Language Models
by: Liu, Alisa, et al.
Published: (2025)
by: Liu, Alisa, et al.
Published: (2025)
Zeroth-Order Optimization at the Edge of Stability
by: Song, Minhak, et al.
Published: (2026)
by: Song, Minhak, et al.
Published: (2026)
Scalable Fingerprinting of Large Language Models
by: Nasery, Anshul, et al.
Published: (2025)
by: Nasery, Anshul, et al.
Published: (2025)
EXaMCaP: Subset Selection with Entropy Gain Maximization for Probing Capability Gains of Large Chart Understanding Training Sets
by: Liu, Jiapeng, et al.
Published: (2026)
by: Liu, Jiapeng, et al.
Published: (2026)
Decoupled-Value Attention for Prior-Data Fitted Networks: GP Inference for Physical Equations
by: Sharma, Kaustubh, et al.
Published: (2025)
by: Sharma, Kaustubh, et al.
Published: (2025)
Zeroth-Order Optimization Finds Flat Minima
by: Zhang, Liang, et al.
Published: (2025)
by: Zhang, Liang, et al.
Published: (2025)
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
One-shot Empirical Privacy Estimation for Federated Learning
by: Andrew, Galen, et al.
Published: (2023)
by: Andrew, Galen, et al.
Published: (2023)
Distill, Forget, Repeat: A Framework for Continual Unlearning in Text-to-Image Diffusion Models
by: George, Naveen, et al.
Published: (2025)
by: George, Naveen, et al.
Published: (2025)
Generative Pre-Training of Time-Series Data for Unsupervised Fault Detection in Semiconductor Manufacturing
by: Lee, Sewoong, et al.
Published: (2023)
by: Lee, Sewoong, et al.
Published: (2023)
Erasure Coded Neural Network Inference via Fisher Averaging
by: Jhunjhunwala, Divyansh, et al.
Published: (2024)
by: Jhunjhunwala, Divyansh, et al.
Published: (2024)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
by: Zhao, Siyan, et al.
Published: (2026)
by: Zhao, Siyan, et al.
Published: (2026)
Can Public Large Language Models Help Private Cross-device Federated Learning?
by: Wang, Boxin, et al.
Published: (2023)
by: Wang, Boxin, et al.
Published: (2023)
Better Alignment with Instruction Back-and-Forth Translation
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
Recycling the Web: A Method to Enhance Pre-training Data Quality and Quantity for Language Models
by: Nguyen, Thao, et al.
Published: (2025)
by: Nguyen, Thao, et al.
Published: (2025)
Self-Distillation Enables Continual Learning
by: Shenfeld, Idan, et al.
Published: (2026)
by: Shenfeld, Idan, et al.
Published: (2026)
Self-Distilled RLVR
by: Yang, Chenxu, et al.
Published: (2026)
by: Yang, Chenxu, et al.
Published: (2026)
A Teacher-Free Graph Knowledge Distillation Framework with Dual Self-Distillation
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
GraSSRep: Graph-Based Self-Supervised Learning for Repeat Detection in Metagenomic Assembly
by: Azizpour, Ali, et al.
Published: (2024)
by: Azizpour, Ali, et al.
Published: (2024)
PickLLM: Context-Aware RL-Assisted Large Language Model Routing
by: Sikeridis, Dimitrios, et al.
Published: (2024)
by: Sikeridis, Dimitrios, et al.
Published: (2024)
Data-Efficient Strategies for Probabilistic Voltage Envelopes under Network Contingencies
by: Pareek, Parikshit, et al.
Published: (2023)
by: Pareek, Parikshit, et al.
Published: (2023)
Small LLMs with Expert Blocks Are Good Enough for Hyperparamter Tuning
by: Naphade, Om, et al.
Published: (2025)
by: Naphade, Om, et al.
Published: (2025)
Learning Power Flow with Confidence: A Probabilistic Guarantee Framework for Voltage Risk
by: Pareek, Parikshit, et al.
Published: (2023)
by: Pareek, Parikshit, et al.
Published: (2023)
Foundation model for mass spectrometry proteomics
by: Sanders, Justin, et al.
Published: (2025)
by: Sanders, Justin, et al.
Published: (2025)
Mechanistic Evidence for Spectral Structures in Prior-Data Fitted Networks
by: Sharma, Kaustubh, et al.
Published: (2026)
by: Sharma, Kaustubh, et al.
Published: (2026)
Similar Items
-
Understanding the Gain from Data Filtering in Multimodal Contrastive Learning
by: Pareek, Divyansh, et al.
Published: (2025) -
PLeaS -- Merging Models with Permutations and Least Squares
by: Nasery, Anshul, et al.
Published: (2024) -
Self Distillation via Iterative Constructive Perturbations
by: Dave, Maheak, et al.
Published: (2025) -
Randomization Techniques to Mitigate the Risk of Copyright Infringement
by: Chen, Wei-Ning, et al.
Published: (2024) -
Training AI to be Loyal
by: Oh, Sewoong, et al.
Published: (2025)