Progressive distillation induces an implicit curriculum
Fuente:
arXiv
Guardado en:
| Autores principales: | Panigrahi, Abhishek, Liu, Bingbin, Malladi, Sadhika, Risteski, Andrej, Goel, Surbhi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
por: Panigrahi, Abhishek, et al.
Publicado: (2025)
por: Panigrahi, Abhishek, et al.
Publicado: (2025)
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
por: Malladi, Sadhika, et al.
Publicado: (2022)
por: Malladi, Sadhika, et al.
Publicado: (2022)
Trainable Transformer in Transformer
por: Panigrahi, Abhishek, et al.
Publicado: (2023)
por: Panigrahi, Abhishek, et al.
Publicado: (2023)
Understanding Augmentation-based Self-Supervised Representation Learning via RKHS Approximation and Regression
por: Zhai, Runtian, et al.
Publicado: (2023)
por: Zhai, Runtian, et al.
Publicado: (2023)
Fit Like You Sample: Sample-Efficient Generalized Score Matching from Fast Mixing Diffusions
por: Qin, Yilong, et al.
Publicado: (2023)
por: Qin, Yilong, et al.
Publicado: (2023)
Provable unlearning in topic modeling and downstream tasks
por: Wei, Stanley, et al.
Publicado: (2024)
por: Wei, Stanley, et al.
Publicado: (2024)
Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws
por: Jiang, Yiding, et al.
Publicado: (2024)
por: Jiang, Yiding, et al.
Publicado: (2024)
The Marginal Value of Momentum for Small Learning Rate SGD
por: Wang, Runzhe, et al.
Publicado: (2023)
por: Wang, Runzhe, et al.
Publicado: (2023)
A computational phase transition for learning-to-sample from Ising models
por: Risteski, Andrej, et al.
Publicado: (2026)
por: Risteski, Andrej, et al.
Publicado: (2026)
The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives
por: Moitra, Ankur, et al.
Publicado: (2026)
por: Moitra, Ankur, et al.
Publicado: (2026)
Steering diffusion models with quadratic rewards: a fine-grained analysis
por: Moitra, Ankur, et al.
Publicado: (2026)
por: Moitra, Ankur, et al.
Publicado: (2026)
Weight Clipping for Robust Conformal Inference under Unbounded Covariate Shifts
por: Wang, James, et al.
Publicado: (2026)
por: Wang, James, et al.
Publicado: (2026)
Reliable Abstention under Adversarial Injections: Tight Lower Bounds and New Upper Bounds
por: Edelman, Ezra, et al.
Publicado: (2026)
por: Edelman, Ezra, et al.
Publicado: (2026)
LESS: Selecting Influential Data for Targeted Instruction Tuning
por: Xia, Mengzhou, et al.
Publicado: (2024)
por: Xia, Mengzhou, et al.
Publicado: (2024)
Tolerant Algorithms for Learning with Arbitrary Covariate Shift
por: Goel, Surbhi, et al.
Publicado: (2024)
por: Goel, Surbhi, et al.
Publicado: (2024)
Adversarial Resilience in Sequential Prediction via Abstention
por: Goel, Surbhi, et al.
Publicado: (2023)
por: Goel, Surbhi, et al.
Publicado: (2023)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
por: Razin, Noam, et al.
Publicado: (2024)
por: Razin, Noam, et al.
Publicado: (2024)
Learning When to Stop: Selective Imitation Learning Under Arbitrary Dynamics Shift
por: Goel, Surbhi, et al.
Publicado: (2026)
por: Goel, Surbhi, et al.
Publicado: (2026)
On the Benefits of Memory for Modeling Time-Dependent PDEs
por: Ruiz, Ricardo Buitrago, et al.
Publicado: (2024)
por: Ruiz, Ricardo Buitrago, et al.
Publicado: (2024)
Complexity Matters: Dynamics of Feature Learning in the Presence of Spurious Correlations
por: Qiu, GuanWen, et al.
Publicado: (2024)
por: Qiu, GuanWen, et al.
Publicado: (2024)
Fine-Tuning Language Models with Just Forward Passes
por: Malladi, Sadhika, et al.
Publicado: (2023)
por: Malladi, Sadhika, et al.
Publicado: (2023)
Taming Imperfect Process Verifiers: A Sampling Perspective on Backtracking
por: Rohatgi, Dhruv, et al.
Publicado: (2025)
por: Rohatgi, Dhruv, et al.
Publicado: (2025)
Preference Learning Algorithms Do Not Learn Preference Rankings
por: Chen, Angelica, et al.
Publicado: (2024)
por: Chen, Angelica, et al.
Publicado: (2024)
Towards characterizing the value of edge embeddings in Graph Neural Networks
por: Rohatgi, Dhruv, et al.
Publicado: (2024)
por: Rohatgi, Dhruv, et al.
Publicado: (2024)
On the Query Complexity of Verifier-Assisted Language Generation
por: Botta, Edoardo, et al.
Publicado: (2025)
por: Botta, Edoardo, et al.
Publicado: (2025)
Explicitly Encoding Structural Symmetry is Key to Length Generalization in Arithmetic Tasks
por: Sabbaghi, Mahdi, et al.
Publicado: (2024)
por: Sabbaghi, Mahdi, et al.
Publicado: (2024)
On the Power of Context-Enhanced Learning in LLMs
por: Zhu, Xingyu, et al.
Publicado: (2025)
por: Zhu, Xingyu, et al.
Publicado: (2025)
Stochastic Bandits with ReLU Neural Networks
por: Xu, Kan, et al.
Publicado: (2024)
por: Xu, Kan, et al.
Publicado: (2024)
Efficient Stagewise Pretraining via Progressive Subnetworks
por: Panigrahi, Abhishek, et al.
Publicado: (2024)
por: Panigrahi, Abhishek, et al.
Publicado: (2024)
Why Do Transformers Fail to Forecast Time Series In-Context?
por: Zhou, Yufa, et al.
Publicado: (2025)
por: Zhou, Yufa, et al.
Publicado: (2025)
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
por: Edelman, Benjamin L., et al.
Publicado: (2024)
por: Edelman, Benjamin L., et al.
Publicado: (2024)
Testing Noise Assumptions of Learning Algorithms
por: Goel, Surbhi, et al.
Publicado: (2025)
por: Goel, Surbhi, et al.
Publicado: (2025)
The Coverage Principle: How Pre-Training Enables Post-Training
por: Chen, Fan, et al.
Publicado: (2025)
por: Chen, Fan, et al.
Publicado: (2025)
Probabilistic Stability Guarantees for Feature Attributions
por: Jin, Helen, et al.
Publicado: (2025)
por: Jin, Helen, et al.
Publicado: (2025)
Tractable Agreement Protocols
por: Collina, Natalie, et al.
Publicado: (2024)
por: Collina, Natalie, et al.
Publicado: (2024)
Representing Rule-based Chatbots with Transformers
por: Friedman, Dan, et al.
Publicado: (2024)
por: Friedman, Dan, et al.
Publicado: (2024)
Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical Guidelines
por: Li, Yuchen, et al.
Publicado: (2024)
por: Li, Yuchen, et al.
Publicado: (2024)
Conformal Language Model Reasoning with Coherent Factuality
por: Rubin-Toles, Maxon, et al.
Publicado: (2025)
por: Rubin-Toles, Maxon, et al.
Publicado: (2025)
Skill-Targeted Adaptive Training
por: He, Yinghui, et al.
Publicado: (2025)
por: He, Yinghui, et al.
Publicado: (2025)
CodePDE: An Inference Framework for LLM-driven PDE Solver Generation
por: Li, Shanda, et al.
Publicado: (2025)
por: Li, Shanda, et al.
Publicado: (2025)
Ejemplares similares
-
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
por: Panigrahi, Abhishek, et al.
Publicado: (2025) -
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
por: Malladi, Sadhika, et al.
Publicado: (2022) -
Trainable Transformer in Transformer
por: Panigrahi, Abhishek, et al.
Publicado: (2023) -
Understanding Augmentation-based Self-Supervised Representation Learning via RKHS Approximation and Regression
por: Zhai, Runtian, et al.
Publicado: (2023) -
Fit Like You Sample: Sample-Efficient Generalized Score Matching from Fast Mixing Diffusions
por: Qin, Yilong, et al.
Publicado: (2023)