Celo: Training Versatile Learned Optimizers on a Compute Diet
Fuente:
arXiv
Saved in:
| Main Authors: | Moudgil, Abhinav, Knyazev, Boris, Lajoie, Guillaume, Belilovsky, Eugene |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Celo2: Towards Learned Optimization Free Lunch
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Accelerating Training with Neuron Interaction and Nowcasting Networks
by: Knyazev, Boris, et al.
Published: (2024)
by: Knyazev, Boris, et al.
Published: (2024)
Meta-learning Optimizers for Communication-Efficient Learning
by: Joseph, Charles-Étienne, et al.
Published: (2023)
by: Joseph, Charles-Étienne, et al.
Published: (2023)
PyLO: Towards Accessible Learned Optimizers in PyTorch
by: Janson, Paul, et al.
Published: (2025)
by: Janson, Paul, et al.
Published: (2025)
$μ$LO: Compute-Efficient Meta-Generalization of Learned Optimizers
by: Thérien, Benjamin, et al.
Published: (2024)
by: Thérien, Benjamin, et al.
Published: (2024)
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
by: Nanfack, Geraldin, et al.
Published: (2025)
by: Nanfack, Geraldin, et al.
Published: (2025)
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
by: Miahi, Erfan, et al.
Published: (2026)
by: Miahi, Erfan, et al.
Published: (2026)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
by: Legate, Gwen, et al.
Published: (2025)
by: Legate, Gwen, et al.
Published: (2025)
Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention
by: Singh, Vaibhav, et al.
Published: (2024)
by: Singh, Vaibhav, et al.
Published: (2024)
Model Breadcrumbs: Scaling Multi-Task Model Merging with Sparse Masks
by: Davari, MohammadReza, et al.
Published: (2023)
by: Davari, MohammadReza, et al.
Published: (2023)
Heterogeneous Low-Bandwidth Pre-Training of LLMs
by: Obeidi, Yazan, et al.
Published: (2026)
by: Obeidi, Yazan, et al.
Published: (2026)
Iterative Amortized Inference: Unifying In-Context Learning and Learned Optimizers
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
Rethinking Prompt Optimization: Reinforcement, Diversification, and Migration in Blackbox LLMs
by: Davari, MohammadReza, et al.
Published: (2025)
by: Davari, MohammadReza, et al.
Published: (2025)
When Data Falls Short: Grokking Below the Critical Threshold
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Stabilizing Native Low-Rank LLM Pretraining
by: Janson, Paul, et al.
Published: (2026)
by: Janson, Paul, et al.
Published: (2026)
PETRA: Parallel End-to-end Training with Reversible Architectures
by: Rivaud, Stéphane, et al.
Published: (2024)
by: Rivaud, Stéphane, et al.
Published: (2024)
Efficient Refusal Ablation in LLM through Optimal Transport
by: Nanfack, Geraldin, et al.
Published: (2026)
by: Nanfack, Geraldin, et al.
Published: (2026)
Communication Efficient LLM Pre-training with SparseLoCo
by: Sarfi, Amir, et al.
Published: (2025)
by: Sarfi, Amir, et al.
Published: (2025)
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
by: Hameed, Humza Wajid, et al.
Published: (2024)
by: Hameed, Humza Wajid, et al.
Published: (2024)
Bidirectional Information Flow (BIF) -- A Sample Efficient Hierarchical Gaussian Process for Bayesian Optimization
by: Guerra, Juan D., et al.
Published: (2025)
by: Guerra, Juan D., et al.
Published: (2025)
Lazy vs hasty: linearization in deep networks impacts learning schedule based on example difficulty
by: George, Thomas, et al.
Published: (2022)
by: George, Thomas, et al.
Published: (2022)
AdaFisher: Adaptive Second Order Optimization via Fisher Information
by: Gomes, Damien Martins, et al.
Published: (2024)
by: Gomes, Damien Martins, et al.
Published: (2024)
DragD3D: Realistic Mesh Editing with Rigidity Control Driven by 2D Diffusion Priors
by: Xie, Tianhao, et al.
Published: (2023)
by: Xie, Tianhao, et al.
Published: (2023)
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
by: Nabli, Adel, et al.
Published: (2024)
by: Nabli, Adel, et al.
Published: (2024)
From Feature Visualization to Visual Circuits: Effect of Adversarial Model Manipulation
by: Nanfack, Geraldin, et al.
Published: (2024)
by: Nanfack, Geraldin, et al.
Published: (2024)
MuLoCo: Muon is a practical inner optimizer for DiLoCo
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
Learning to Rank with Variable Result Presentation Lengths
by: Knyazev, Norman, et al.
Published: (2025)
by: Knyazev, Norman, et al.
Published: (2025)
Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks
by: Williams, Ezekiel, et al.
Published: (2026)
by: Williams, Ezekiel, et al.
Published: (2026)
Beyond Distribution Sharpening: The Importance of Task Rewards
by: Mittal, Sarthak, et al.
Published: (2026)
by: Mittal, Sarthak, et al.
Published: (2026)
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Generating $π$-Functional Molecules Using STGG+ with Active Learning
by: Jolicoeur-Martineau, Alexia, et al.
Published: (2025)
by: Jolicoeur-Martineau, Alexia, et al.
Published: (2025)
LoGAH: Predicting 774-Million-Parameter Transformers using Graph HyperNetworks with 1/100 Parameters
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
WASH: Train your Ensemble with Communication-Efficient Weight Shuffling, then Average
by: Fournier, Louis, et al.
Published: (2024)
by: Fournier, Louis, et al.
Published: (2024)
Model Parallelism With Subnetwork Data Parallelism
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Non-Uniform Parameter-Wise Model Merging
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
Incentivizing Permissionless Distributed Learning of LLMs
by: Lidin, Joel, et al.
Published: (2025)
by: Lidin, Joel, et al.
Published: (2025)
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
by: Huang, Ningyuan, et al.
Published: (2025)
by: Huang, Ningyuan, et al.
Published: (2025)
Attention to Mamba: A Recipe for Cross-Architecture Distillation
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
In-Context Parametric Inference: Point or Distribution Estimators?
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
Similar Items
-
Celo2: Towards Learned Optimization Free Lunch
by: Moudgil, Abhinav, et al.
Published: (2026) -
Accelerating Training with Neuron Interaction and Nowcasting Networks
by: Knyazev, Boris, et al.
Published: (2024) -
Meta-learning Optimizers for Communication-Efficient Learning
by: Joseph, Charles-Étienne, et al.
Published: (2023) -
PyLO: Towards Accessible Learned Optimizers in PyTorch
by: Janson, Paul, et al.
Published: (2025) -
$μ$LO: Compute-Efficient Meta-Generalization of Learned Optimizers
by: Thérien, Benjamin, et al.
Published: (2024)