Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Topollai, Kristi, Ma, Allan, Dimlioglu, Tolga, Tay, Sui Jiet, Choromanska, Anna |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Worker Disagreement Reveals Sharp Directions in Local SGD
by: Dimlioglu, Tolga, et al.
Published: (2026)
by: Dimlioglu, Tolga, et al.
Published: (2026)
Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
by: Topollai, Kristi, et al.
Published: (2025)
by: Topollai, Kristi, et al.
Published: (2025)
Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets
by: Topollai, Kristi, et al.
Published: (2026)
by: Topollai, Kristi, et al.
Published: (2026)
Task-Level Contrastiveness for Cross-Domain Few-Shot Learning
by: Topollai, Kristi, et al.
Published: (2025)
by: Topollai, Kristi, et al.
Published: (2025)
Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
by: Dimlioglu, Tolga, et al.
Published: (2025)
by: Dimlioglu, Tolga, et al.
Published: (2025)
GRAWA: Gradient-based Weighted Averaging for Distributed Training of Deep Learning Models
by: Dimlioglu, Tolga, et al.
Published: (2024)
by: Dimlioglu, Tolga, et al.
Published: (2024)
Streamlining Industrial Contract Management with Retrieval-Augmented LLMs
by: Topollai, Kristi, et al.
Published: (2025)
by: Topollai, Kristi, et al.
Published: (2025)
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
by: Hemadri, Raghu Vamshi, et al.
Published: (2025)
by: Hemadri, Raghu Vamshi, et al.
Published: (2025)
Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object Detection
by: Zhu, Haoran, et al.
Published: (2025)
by: Zhu, Haoran, et al.
Published: (2025)
Multimodal Data Curation via Object Detection and Filter Ensembles
by: Huang, Tzu-Heng, et al.
Published: (2024)
by: Huang, Tzu-Heng, et al.
Published: (2024)
A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Pearls from Pebbles: Improved Confidence Functions for Auto-labeling
by: Vishwakarma, Harit, et al.
Published: (2024)
by: Vishwakarma, Harit, et al.
Published: (2024)
Adjacent Leader Decentralized Stochastic Gradient Descent
by: He, Haoze, et al.
Published: (2024)
by: He, Haoze, et al.
Published: (2024)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems
by: Dimlioglu, Tolga, et al.
Published: (2026)
by: Dimlioglu, Tolga, et al.
Published: (2026)
TAME: Task Agnostic Continual Learning using Multiple Experts
by: Zhu, Haoran, et al.
Published: (2022)
by: Zhu, Haoran, et al.
Published: (2022)
An Invariant Information Geometric Method for High-Dimensional Online Optimization
by: Zhang, Zhengfei, et al.
Published: (2024)
by: Zhang, Zhengfei, et al.
Published: (2024)
Safe Bayesian Optimization for the Control of High-Dimensional Embodied Systems
by: Wei, Yunyue, et al.
Published: (2024)
by: Wei, Yunyue, et al.
Published: (2024)
SNOO: Step-K Nesterov Outer Optimizer - The Surprising Effectiveness of Nesterov Momentum Applied to Pseudo-Gradients
by: Kallusky, Dominik, et al.
Published: (2025)
by: Kallusky, Dominik, et al.
Published: (2025)
Zero-Shot Cross-City Generalization in End-to-End Autonomous Driving: Self-Supervised versus Supervised Representations
by: Naeinian, Fatemeh, et al.
Published: (2026)
by: Naeinian, Fatemeh, et al.
Published: (2026)
Restarted contractive operators to learn at equilibrium
by: Davy, Leo, et al.
Published: (2025)
by: Davy, Leo, et al.
Published: (2025)
Scalable Exploration for High-Dimensional Continuous Control via Value-Guided Flow
by: Wei, Yunyue, et al.
Published: (2026)
by: Wei, Yunyue, et al.
Published: (2026)
Feature Clock: High-Dimensional Effects in Two-Dimensional Plots
by: Ovcharenko, Olga, et al.
Published: (2024)
by: Ovcharenko, Olga, et al.
Published: (2024)
Efficient Restarts in Non-Stationary Model-Free Reinforcement Learning
by: Nonaka, Hiroshi, et al.
Published: (2025)
by: Nonaka, Hiroshi, et al.
Published: (2025)
Understanding High-Dimensional Bayesian Optimization
by: Papenmeier, Leonard, et al.
Published: (2025)
by: Papenmeier, Leonard, et al.
Published: (2025)
Solving Diffusion Inverse Problems with Restart Posterior Sampling
by: Ahmed, Bilal, et al.
Published: (2025)
by: Ahmed, Bilal, et al.
Published: (2025)
Data Scaling Laws for End-to-End Autonomous Driving
by: Naumann, Alexander, et al.
Published: (2025)
by: Naumann, Alexander, et al.
Published: (2025)
Convex Formulations for Training Two-Layer ReLU Neural Networks
by: Prakhya, Karthik, et al.
Published: (2024)
by: Prakhya, Karthik, et al.
Published: (2024)
FISMO: Fisher-Structured Momentum-Orthogonalized Optimizer
by: Xu, Chenrui, et al.
Published: (2026)
by: Xu, Chenrui, et al.
Published: (2026)
Efficiency of Parallel and Restart Exploration Strategies in Model Free Stochastic Simulations
by: Garcia, Ernesto, et al.
Published: (2025)
by: Garcia, Ernesto, et al.
Published: (2025)
Bridging Training and Merging Through Momentum-Aware Optimization
by: Moayedikia, Alireza, et al.
Published: (2025)
by: Moayedikia, Alireza, et al.
Published: (2025)
HOME-3: High-Order Momentum Estimator with Third-Power Gradient for Convex and Smooth Nonconvex Optimization
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
Stochastic Difference-of-Convex Optimization with Momentum
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
DeMo: Decoupled Momentum Optimization
by: Peng, Bowen, et al.
Published: (2024)
by: Peng, Bowen, et al.
Published: (2024)
On the Limits of Momentum in Decentralized and Federated Optimization
by: Zaccone, Riccardo, et al.
Published: (2025)
by: Zaccone, Riccardo, et al.
Published: (2025)
Adaptive Linear Embedding for Nonstationary High-Dimensional Optimization
by: Wen, Yuejiang, et al.
Published: (2025)
by: Wen, Yuejiang, et al.
Published: (2025)
An Adaptive Dropout Approach for High-Dimensional Bayesian Optimization
by: Huang, Jundi, et al.
Published: (2025)
by: Huang, Jundi, et al.
Published: (2025)
Learning When to Restart: Nonstationary Newsvendor from Uncensored to Censored Demand
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
Optimizing the Adversarial Perturbation with a Momentum-based Adaptive Matrix
by: Tao, Wei, et al.
Published: (2025)
by: Tao, Wei, et al.
Published: (2025)
Prescriptive PCA: Dimensionality Reduction for Two-stage Stochastic Optimization
by: He, Long, et al.
Published: (2023)
by: He, Long, et al.
Published: (2023)
Similar Items
-
Worker Disagreement Reveals Sharp Directions in Local SGD
by: Dimlioglu, Tolga, et al.
Published: (2026) -
Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
by: Topollai, Kristi, et al.
Published: (2025) -
Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets
by: Topollai, Kristi, et al.
Published: (2026) -
Task-Level Contrastiveness for Cross-Domain Few-Shot Learning
by: Topollai, Kristi, et al.
Published: (2025) -
Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
by: Dimlioglu, Tolga, et al.
Published: (2025)