On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
Fuente:
arXiv
Saved in:
| Main Authors: | Malladi, Sadhika, Lyu, Kaifeng, Panigrahi, Abhishek, Arora, Sanjeev |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
Provable unlearning in topic modeling and downstream tasks
by: Wei, Stanley, et al.
Published: (2024)
by: Wei, Stanley, et al.
Published: (2024)
Progressive distillation induces an implicit curriculum
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
by: Panigrahi, Abhishek, et al.
Published: (2025)
by: Panigrahi, Abhishek, et al.
Published: (2025)
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)
by: Wang, Runzhe, et al.
Published: (2023)
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024)
by: Xia, Mengzhou, et al.
Published: (2024)
Skill-Targeted Adaptive Training
by: He, Yinghui, et al.
Published: (2025)
by: He, Yinghui, et al.
Published: (2025)
On the Power of Context-Enhanced Learning in LLMs
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
A Quadratic Synchronization Rule for Distributed Deep Learning
by: Gu, Xinran, et al.
Published: (2023)
by: Gu, Xinran, et al.
Published: (2023)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws
by: Jiang, Yiding, et al.
Published: (2024)
by: Jiang, Yiding, et al.
Published: (2024)
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)
by: Malladi, Sadhika, et al.
Published: (2023)
Weak-to-Strong Generalization Even in Random Feature Networks, Provably
by: Medvedev, Marko, et al.
Published: (2025)
by: Medvedev, Marko, et al.
Published: (2025)
Representing Rule-based Chatbots with Transformers
by: Friedman, Dan, et al.
Published: (2024)
by: Friedman, Dan, et al.
Published: (2024)
Preference Learning Algorithms Do Not Learn Preference Rankings
by: Chen, Angelica, et al.
Published: (2024)
by: Chen, Angelica, et al.
Published: (2024)
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
by: Lyu, Kaifeng, et al.
Published: (2024)
by: Lyu, Kaifeng, et al.
Published: (2024)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
by: Park, Simon, et al.
Published: (2025)
by: Park, Simon, et al.
Published: (2025)
Efficient Stagewise Pretraining via Progressive Subnetworks
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks
by: Li, Binghui, et al.
Published: (2024)
by: Li, Binghui, et al.
Published: (2024)
Unrealized Expectations: Comparing AI Methods vs Classical Algorithms for Maximum Independent Set
by: Wu, Yikai, et al.
Published: (2025)
by: Wu, Yikai, et al.
Published: (2025)
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
by: Shang, Shuning, et al.
Published: (2026)
by: Shang, Shuning, et al.
Published: (2026)
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
by: He, Yinghui, et al.
Published: (2025)
by: He, Yinghui, et al.
Published: (2025)
How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
by: Park, Simon, et al.
Published: (2025)
by: Park, Simon, et al.
Published: (2025)
Adam Reduces a Unique Form of Sharpness: Theoretical Insights Near the Minimizer Manifold
by: Li, Xinghan, et al.
Published: (2025)
by: Li, Xinghan, et al.
Published: (2025)
The Coverage Principle: How Pre-Training Enables Post-Training
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear Regression
by: Yan, Tingkai, et al.
Published: (2025)
by: Yan, Tingkai, et al.
Published: (2025)
AI-Assisted Generation of Difficult Math Questions
by: Shah, Vedant, et al.
Published: (2024)
by: Shah, Vedant, et al.
Published: (2024)
On the Impossibility of Retrain Equivalence in Machine Unlearning
by: Yu, Jiatong, et al.
Published: (2025)
by: Yu, Jiatong, et al.
Published: (2025)
RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
by: Wen, Kaiyue, et al.
Published: (2024)
by: Wen, Kaiyue, et al.
Published: (2024)
Gradient-flow SDEs have unique transient population dynamics
by: Guan, Vincent, et al.
Published: (2025)
by: Guan, Vincent, et al.
Published: (2025)
Adaptive Regularization for Large-Scale Sparse Feature Embedding Models
by: Li, Mang, et al.
Published: (2025)
by: Li, Mang, et al.
Published: (2025)
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
Shift is Good: Mismatched Data Mixing Improves Test Performance
by: Medvedev, Marko, et al.
Published: (2025)
by: Medvedev, Marko, et al.
Published: (2025)
Impact of Feature Scaling on the Performance of Classification Algorithms: a Comparative Study on Wine Dataset
by: Arora, Suhani
Published: (2026)
by: Arora, Suhani
Published: (2026)
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
by: Kaur, Simran, et al.
Published: (2024)
by: Kaur, Simran, et al.
Published: (2024)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
Distance-Aware Muon: Adaptive Step Scaling for Normalized Optimization
by: Demidovich, Yury, et al.
Published: (2026)
by: Demidovich, Yury, et al.
Published: (2026)
Gradient Scaling Effects in Adaptive Spectral PINNs for Stiff Nonlinear ODEs
by: Yepes, Isabela M., et al.
Published: (2026)
by: Yepes, Isabela M., et al.
Published: (2026)
Analysis of Estimating the Bayes Rule for Gaussian Mixture Models with a Specified Missing-Data Mechanism
by: Lyu, Ziyang
Published: (2022)
by: Lyu, Ziyang
Published: (2022)
Orthogonal Gradient Boosting for Simpler Additive Rule Ensembles
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
Similar Items
-
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023) -
Provable unlearning in topic modeling and downstream tasks
by: Wei, Stanley, et al.
Published: (2024) -
Progressive distillation induces an implicit curriculum
by: Panigrahi, Abhishek, et al.
Published: (2024) -
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
by: Panigrahi, Abhishek, et al.
Published: (2025) -
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)