Progressive distillation induces an implicit curriculum
Fuente:
arXiv
Saved in:
| Main Authors: | Panigrahi, Abhishek, Liu, Bingbin, Malladi, Sadhika, Risteski, Andrej, Goel, Surbhi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
by: Panigrahi, Abhishek, et al.
Published: (2025)
by: Panigrahi, Abhishek, et al.
Published: (2025)
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
by: Malladi, Sadhika, et al.
Published: (2022)
by: Malladi, Sadhika, et al.
Published: (2022)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
Understanding Augmentation-based Self-Supervised Representation Learning via RKHS Approximation and Regression
by: Zhai, Runtian, et al.
Published: (2023)
by: Zhai, Runtian, et al.
Published: (2023)
Fit Like You Sample: Sample-Efficient Generalized Score Matching from Fast Mixing Diffusions
by: Qin, Yilong, et al.
Published: (2023)
by: Qin, Yilong, et al.
Published: (2023)
Provable unlearning in topic modeling and downstream tasks
by: Wei, Stanley, et al.
Published: (2024)
by: Wei, Stanley, et al.
Published: (2024)
Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws
by: Jiang, Yiding, et al.
Published: (2024)
by: Jiang, Yiding, et al.
Published: (2024)
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)
by: Wang, Runzhe, et al.
Published: (2023)
A computational phase transition for learning-to-sample from Ising models
by: Risteski, Andrej, et al.
Published: (2026)
by: Risteski, Andrej, et al.
Published: (2026)
The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives
by: Moitra, Ankur, et al.
Published: (2026)
by: Moitra, Ankur, et al.
Published: (2026)
Steering diffusion models with quadratic rewards: a fine-grained analysis
by: Moitra, Ankur, et al.
Published: (2026)
by: Moitra, Ankur, et al.
Published: (2026)
Weight Clipping for Robust Conformal Inference under Unbounded Covariate Shifts
by: Wang, James, et al.
Published: (2026)
by: Wang, James, et al.
Published: (2026)
Reliable Abstention under Adversarial Injections: Tight Lower Bounds and New Upper Bounds
by: Edelman, Ezra, et al.
Published: (2026)
by: Edelman, Ezra, et al.
Published: (2026)
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024)
by: Xia, Mengzhou, et al.
Published: (2024)
Tolerant Algorithms for Learning with Arbitrary Covariate Shift
by: Goel, Surbhi, et al.
Published: (2024)
by: Goel, Surbhi, et al.
Published: (2024)
Adversarial Resilience in Sequential Prediction via Abstention
by: Goel, Surbhi, et al.
Published: (2023)
by: Goel, Surbhi, et al.
Published: (2023)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
Learning When to Stop: Selective Imitation Learning Under Arbitrary Dynamics Shift
by: Goel, Surbhi, et al.
Published: (2026)
by: Goel, Surbhi, et al.
Published: (2026)
On the Benefits of Memory for Modeling Time-Dependent PDEs
by: Ruiz, Ricardo Buitrago, et al.
Published: (2024)
by: Ruiz, Ricardo Buitrago, et al.
Published: (2024)
Complexity Matters: Dynamics of Feature Learning in the Presence of Spurious Correlations
by: Qiu, GuanWen, et al.
Published: (2024)
by: Qiu, GuanWen, et al.
Published: (2024)
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)
by: Malladi, Sadhika, et al.
Published: (2023)
Taming Imperfect Process Verifiers: A Sampling Perspective on Backtracking
by: Rohatgi, Dhruv, et al.
Published: (2025)
by: Rohatgi, Dhruv, et al.
Published: (2025)
Preference Learning Algorithms Do Not Learn Preference Rankings
by: Chen, Angelica, et al.
Published: (2024)
by: Chen, Angelica, et al.
Published: (2024)
Towards characterizing the value of edge embeddings in Graph Neural Networks
by: Rohatgi, Dhruv, et al.
Published: (2024)
by: Rohatgi, Dhruv, et al.
Published: (2024)
On the Query Complexity of Verifier-Assisted Language Generation
by: Botta, Edoardo, et al.
Published: (2025)
by: Botta, Edoardo, et al.
Published: (2025)
Explicitly Encoding Structural Symmetry is Key to Length Generalization in Arithmetic Tasks
by: Sabbaghi, Mahdi, et al.
Published: (2024)
by: Sabbaghi, Mahdi, et al.
Published: (2024)
On the Power of Context-Enhanced Learning in LLMs
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
Stochastic Bandits with ReLU Neural Networks
by: Xu, Kan, et al.
Published: (2024)
by: Xu, Kan, et al.
Published: (2024)
Efficient Stagewise Pretraining via Progressive Subnetworks
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
Why Do Transformers Fail to Forecast Time Series In-Context?
by: Zhou, Yufa, et al.
Published: (2025)
by: Zhou, Yufa, et al.
Published: (2025)
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
by: Edelman, Benjamin L., et al.
Published: (2024)
by: Edelman, Benjamin L., et al.
Published: (2024)
Testing Noise Assumptions of Learning Algorithms
by: Goel, Surbhi, et al.
Published: (2025)
by: Goel, Surbhi, et al.
Published: (2025)
The Coverage Principle: How Pre-Training Enables Post-Training
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Probabilistic Stability Guarantees for Feature Attributions
by: Jin, Helen, et al.
Published: (2025)
by: Jin, Helen, et al.
Published: (2025)
Tractable Agreement Protocols
by: Collina, Natalie, et al.
Published: (2024)
by: Collina, Natalie, et al.
Published: (2024)
Representing Rule-based Chatbots with Transformers
by: Friedman, Dan, et al.
Published: (2024)
by: Friedman, Dan, et al.
Published: (2024)
Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical Guidelines
by: Li, Yuchen, et al.
Published: (2024)
by: Li, Yuchen, et al.
Published: (2024)
Conformal Language Model Reasoning with Coherent Factuality
by: Rubin-Toles, Maxon, et al.
Published: (2025)
by: Rubin-Toles, Maxon, et al.
Published: (2025)
Skill-Targeted Adaptive Training
by: He, Yinghui, et al.
Published: (2025)
by: He, Yinghui, et al.
Published: (2025)
CodePDE: An Inference Framework for LLM-driven PDE Solver Generation
by: Li, Shanda, et al.
Published: (2025)
by: Li, Shanda, et al.
Published: (2025)
Similar Items
-
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
by: Panigrahi, Abhishek, et al.
Published: (2025) -
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
by: Malladi, Sadhika, et al.
Published: (2022) -
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023) -
Understanding Augmentation-based Self-Supervised Representation Learning via RKHS Approximation and Regression
by: Zhai, Runtian, et al.
Published: (2023) -
Fit Like You Sample: Sample-Efficient Generalized Score Matching from Fast Mixing Diffusions
by: Qin, Yilong, et al.
Published: (2023)