Prodigy: An Expeditiously Adaptive Parameter-Free Learner
Fuente:
arXiv
Saved in:
| Main Authors: | Mishchenko, Konstantin, Defazio, Aaron |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Road Less Scheduled
by: Defazio, Aaron, et al.
Published: (2024)
by: Defazio, Aaron, et al.
Published: (2024)
DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent Method
by: Khaled, Ahmed, et al.
Published: (2023)
by: Khaled, Ahmed, et al.
Published: (2023)
Adaptive Proximal Gradient Method for Convex Optimization
by: Malitsky, Yura, et al.
Published: (2023)
by: Malitsky, Yura, et al.
Published: (2023)
Optimal Linear Decay Learning Rate Schedules and Further Refinements
by: Defazio, Aaron, et al.
Published: (2023)
by: Defazio, Aaron, et al.
Published: (2023)
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Optimal Transport for Machine Learners
by: Peyré, Gabriel
Published: (2025)
by: Peyré, Gabriel
Published: (2025)
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
by: Defazio, Aaron
Published: (2026)
by: Defazio, Aaron
Published: (2026)
A multiobjective continuation method to compute the regularization path of deep neural networks
by: Amakor, Augustina C., et al.
Published: (2023)
by: Amakor, Augustina C., et al.
Published: (2023)
Federated Majorize-Minimization: Beyond Parameter Aggregation
by: Dieuleveut, Aymeric, et al.
Published: (2025)
by: Dieuleveut, Aymeric, et al.
Published: (2025)
Lagrangian Index Policy for Restless Bandits with Average Reward
by: Avrachenkov, Konstantin, et al.
Published: (2024)
by: Avrachenkov, Konstantin, et al.
Published: (2024)
Client-Centric Federated Adaptive Optimization
by: Sun, Jianhui, et al.
Published: (2025)
by: Sun, Jianhui, et al.
Published: (2025)
EXAdam: The Power of Adaptive Cross-Moments
by: Adly, Ahmed M.
Published: (2024)
by: Adly, Ahmed M.
Published: (2024)
Dynamic Memory Based Adaptive Optimization
by: Szegedy, Balázs, et al.
Published: (2024)
by: Szegedy, Balázs, et al.
Published: (2024)
Anytime Training with Schedule-Free Spectral Optimization
by: Apte, Anuj, et al.
Published: (2026)
by: Apte, Anuj, et al.
Published: (2026)
Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cramér Surrogate
by: C., Simo Alami, et al.
Published: (2025)
by: C., Simo Alami, et al.
Published: (2025)
Adaptive Primal-Dual Method for Safe Reinforcement Learning
by: Chen, Weiqin, et al.
Published: (2024)
by: Chen, Weiqin, et al.
Published: (2024)
BAGEL: Projection-Free Algorithm for Adversarially Constrained Online Convex Optimization
by: Lu, Yiyang, et al.
Published: (2025)
by: Lu, Yiyang, et al.
Published: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025)
by: Chen, Zixi, et al.
Published: (2025)
Conformal Prediction-Driven Adaptive Sampling for Digital Water Twins
by: Homaei, Mohammadhossein, et al.
Published: (2025)
by: Homaei, Mohammadhossein, et al.
Published: (2025)
Hidden Convexity of Fair PCA and Fast Solver via Eigenvalue Optimization
by: Shen, Junhui, et al.
Published: (2025)
by: Shen, Junhui, et al.
Published: (2025)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models
by: Zhao, Pengxiang, et al.
Published: (2025)
by: Zhao, Pengxiang, et al.
Published: (2025)
Hindsight-Guided Momentum (HGM) Optimizer: An Approach to Adaptive Learning Rate
by: Sarkar, Krisanu
Published: (2025)
by: Sarkar, Krisanu
Published: (2025)
Kernel-Free Universum Quadratic Surface Twin Support Vector Machines for Imbalanced Data
by: Moosaei, Hossein, et al.
Published: (2024)
by: Moosaei, Hossein, et al.
Published: (2024)
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
by: Defazio, Aaron, et al.
Published: (2025)
by: Defazio, Aaron, et al.
Published: (2025)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
by: Farhat, Yehya, et al.
Published: (2023)
by: Farhat, Yehya, et al.
Published: (2023)
A Methodology Establishing Linear Convergence of Adaptive Gradient Methods under PL Inequality
by: Chakrabarti, Kushal, et al.
Published: (2024)
by: Chakrabarti, Kushal, et al.
Published: (2024)
Balans: Multi-Armed Bandits-based Adaptive Large Neighborhood Search for Mixed-Integer Programming Problem
by: Cai, Junyang, et al.
Published: (2024)
by: Cai, Junyang, et al.
Published: (2024)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
by: Han, X. Y., et al.
Published: (2025)
by: Han, X. Y., et al.
Published: (2025)
Optimal Control Operator Perspective and a Neural Adaptive Spectral Method
by: Feng, Mingquan, et al.
Published: (2024)
by: Feng, Mingquan, et al.
Published: (2024)
Lyapunov Function Consistent Adaptive Network Signal Control with Back Pressure and Reinforcement Learning
by: Ma, Chaolun, et al.
Published: (2022)
by: Ma, Chaolun, et al.
Published: (2022)
Capabilities of Large Language Models in Control Engineering: A Benchmark Study on GPT-4, Claude 3 Opus, and Gemini 1.0 Ultra
by: Kevian, Darioush, et al.
Published: (2024)
by: Kevian, Darioush, et al.
Published: (2024)
PARQ: Piecewise-Affine Regularized Quantization
by: Jin, Lisa, et al.
Published: (2025)
by: Jin, Lisa, et al.
Published: (2025)
DASA: Delay-Adaptive Multi-Agent Stochastic Approximation
by: Fabbro, Nicolò Dal, et al.
Published: (2024)
by: Fabbro, Nicolò Dal, et al.
Published: (2024)
Adaptive Smooth Tchebycheff Attention for Multi-Objective Policy Optimization
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2026)
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2026)
Probabilistic Geometric Alignment via Bayesian Latent Transport for Domain-Adaptive Foundation Models
by: Aueawatthanaphisut, Aueaphum, et al.
Published: (2026)
by: Aueawatthanaphisut, Aueaphum, et al.
Published: (2026)
Why Gradients Rapidly Increase Near the End of Training
by: Defazio, Aaron
Published: (2025)
by: Defazio, Aaron
Published: (2025)
Towards Simple and Provable Parameter-Free Adaptive Gradient Methods
by: Tao, Yuanzhe, et al.
Published: (2024)
by: Tao, Yuanzhe, et al.
Published: (2024)
Optimism Stabilizes Thompson Sampling for Adaptive Inference
by: Yan, Shunxing, et al.
Published: (2026)
by: Yan, Shunxing, et al.
Published: (2026)
Similar Items
-
The Road Less Scheduled
by: Defazio, Aaron, et al.
Published: (2024) -
DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent Method
by: Khaled, Ahmed, et al.
Published: (2023) -
Adaptive Proximal Gradient Method for Convex Optimization
by: Malitsky, Yura, et al.
Published: (2023) -
Optimal Linear Decay Learning Rate Schedules and Further Refinements
by: Defazio, Aaron, et al.
Published: (2023) -
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
by: Mishkin, Aaron, et al.
Published: (2024)