Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Topollai, Kristi, Ma, Allan, Dimlioglu, Tolga, Tay, Sui Jiet, Choromanska, Anna |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Worker Disagreement Reveals Sharp Directions in Local SGD
por: Dimlioglu, Tolga, et al.
Publicado: (2026)
por: Dimlioglu, Tolga, et al.
Publicado: (2026)
Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
por: Topollai, Kristi, et al.
Publicado: (2025)
por: Topollai, Kristi, et al.
Publicado: (2025)
Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets
por: Topollai, Kristi, et al.
Publicado: (2026)
por: Topollai, Kristi, et al.
Publicado: (2026)
Task-Level Contrastiveness for Cross-Domain Few-Shot Learning
por: Topollai, Kristi, et al.
Publicado: (2025)
por: Topollai, Kristi, et al.
Publicado: (2025)
Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
por: Dimlioglu, Tolga, et al.
Publicado: (2025)
por: Dimlioglu, Tolga, et al.
Publicado: (2025)
GRAWA: Gradient-based Weighted Averaging for Distributed Training of Deep Learning Models
por: Dimlioglu, Tolga, et al.
Publicado: (2024)
por: Dimlioglu, Tolga, et al.
Publicado: (2024)
Streamlining Industrial Contract Management with Retrieval-Augmented LLMs
por: Topollai, Kristi, et al.
Publicado: (2025)
por: Topollai, Kristi, et al.
Publicado: (2025)
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
por: Hemadri, Raghu Vamshi, et al.
Publicado: (2025)
por: Hemadri, Raghu Vamshi, et al.
Publicado: (2025)
Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object Detection
por: Zhu, Haoran, et al.
Publicado: (2025)
por: Zhu, Haoran, et al.
Publicado: (2025)
Multimodal Data Curation via Object Detection and Filter Ensembles
por: Huang, Tzu-Heng, et al.
Publicado: (2024)
por: Huang, Tzu-Heng, et al.
Publicado: (2024)
A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization
por: Wang, Jing, et al.
Publicado: (2025)
por: Wang, Jing, et al.
Publicado: (2025)
Pearls from Pebbles: Improved Confidence Functions for Auto-labeling
por: Vishwakarma, Harit, et al.
Publicado: (2024)
por: Vishwakarma, Harit, et al.
Publicado: (2024)
Adjacent Leader Decentralized Stochastic Gradient Descent
por: He, Haoze, et al.
Publicado: (2024)
por: He, Haoze, et al.
Publicado: (2024)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
por: Khaled, Ahmed, et al.
Publicado: (2025)
por: Khaled, Ahmed, et al.
Publicado: (2025)
Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems
por: Dimlioglu, Tolga, et al.
Publicado: (2026)
por: Dimlioglu, Tolga, et al.
Publicado: (2026)
TAME: Task Agnostic Continual Learning using Multiple Experts
por: Zhu, Haoran, et al.
Publicado: (2022)
por: Zhu, Haoran, et al.
Publicado: (2022)
An Invariant Information Geometric Method for High-Dimensional Online Optimization
por: Zhang, Zhengfei, et al.
Publicado: (2024)
por: Zhang, Zhengfei, et al.
Publicado: (2024)
Safe Bayesian Optimization for the Control of High-Dimensional Embodied Systems
por: Wei, Yunyue, et al.
Publicado: (2024)
por: Wei, Yunyue, et al.
Publicado: (2024)
SNOO: Step-K Nesterov Outer Optimizer - The Surprising Effectiveness of Nesterov Momentum Applied to Pseudo-Gradients
por: Kallusky, Dominik, et al.
Publicado: (2025)
por: Kallusky, Dominik, et al.
Publicado: (2025)
Zero-Shot Cross-City Generalization in End-to-End Autonomous Driving: Self-Supervised versus Supervised Representations
por: Naeinian, Fatemeh, et al.
Publicado: (2026)
por: Naeinian, Fatemeh, et al.
Publicado: (2026)
Restarted contractive operators to learn at equilibrium
por: Davy, Leo, et al.
Publicado: (2025)
por: Davy, Leo, et al.
Publicado: (2025)
Scalable Exploration for High-Dimensional Continuous Control via Value-Guided Flow
por: Wei, Yunyue, et al.
Publicado: (2026)
por: Wei, Yunyue, et al.
Publicado: (2026)
Feature Clock: High-Dimensional Effects in Two-Dimensional Plots
por: Ovcharenko, Olga, et al.
Publicado: (2024)
por: Ovcharenko, Olga, et al.
Publicado: (2024)
Efficient Restarts in Non-Stationary Model-Free Reinforcement Learning
por: Nonaka, Hiroshi, et al.
Publicado: (2025)
por: Nonaka, Hiroshi, et al.
Publicado: (2025)
Understanding High-Dimensional Bayesian Optimization
por: Papenmeier, Leonard, et al.
Publicado: (2025)
por: Papenmeier, Leonard, et al.
Publicado: (2025)
Solving Diffusion Inverse Problems with Restart Posterior Sampling
por: Ahmed, Bilal, et al.
Publicado: (2025)
por: Ahmed, Bilal, et al.
Publicado: (2025)
Data Scaling Laws for End-to-End Autonomous Driving
por: Naumann, Alexander, et al.
Publicado: (2025)
por: Naumann, Alexander, et al.
Publicado: (2025)
Convex Formulations for Training Two-Layer ReLU Neural Networks
por: Prakhya, Karthik, et al.
Publicado: (2024)
por: Prakhya, Karthik, et al.
Publicado: (2024)
FISMO: Fisher-Structured Momentum-Orthogonalized Optimizer
por: Xu, Chenrui, et al.
Publicado: (2026)
por: Xu, Chenrui, et al.
Publicado: (2026)
Efficiency of Parallel and Restart Exploration Strategies in Model Free Stochastic Simulations
por: Garcia, Ernesto, et al.
Publicado: (2025)
por: Garcia, Ernesto, et al.
Publicado: (2025)
Bridging Training and Merging Through Momentum-Aware Optimization
por: Moayedikia, Alireza, et al.
Publicado: (2025)
por: Moayedikia, Alireza, et al.
Publicado: (2025)
HOME-3: High-Order Momentum Estimator with Third-Power Gradient for Convex and Smooth Nonconvex Optimization
por: Zhang, Wei, et al.
Publicado: (2025)
por: Zhang, Wei, et al.
Publicado: (2025)
Stochastic Difference-of-Convex Optimization with Momentum
por: Chayti, El Mahdi, et al.
Publicado: (2025)
por: Chayti, El Mahdi, et al.
Publicado: (2025)
DeMo: Decoupled Momentum Optimization
por: Peng, Bowen, et al.
Publicado: (2024)
por: Peng, Bowen, et al.
Publicado: (2024)
On the Limits of Momentum in Decentralized and Federated Optimization
por: Zaccone, Riccardo, et al.
Publicado: (2025)
por: Zaccone, Riccardo, et al.
Publicado: (2025)
Adaptive Linear Embedding for Nonstationary High-Dimensional Optimization
por: Wen, Yuejiang, et al.
Publicado: (2025)
por: Wen, Yuejiang, et al.
Publicado: (2025)
An Adaptive Dropout Approach for High-Dimensional Bayesian Optimization
por: Huang, Jundi, et al.
Publicado: (2025)
por: Huang, Jundi, et al.
Publicado: (2025)
Learning When to Restart: Nonstationary Newsvendor from Uncensored to Censored Demand
por: Chen, Xin, et al.
Publicado: (2025)
por: Chen, Xin, et al.
Publicado: (2025)
Optimizing the Adversarial Perturbation with a Momentum-based Adaptive Matrix
por: Tao, Wei, et al.
Publicado: (2025)
por: Tao, Wei, et al.
Publicado: (2025)
Prescriptive PCA: Dimensionality Reduction for Two-stage Stochastic Optimization
por: He, Long, et al.
Publicado: (2023)
por: He, Long, et al.
Publicado: (2023)
Ejemplares similares
-
Worker Disagreement Reveals Sharp Directions in Local SGD
por: Dimlioglu, Tolga, et al.
Publicado: (2026) -
Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
por: Topollai, Kristi, et al.
Publicado: (2025) -
Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets
por: Topollai, Kristi, et al.
Publicado: (2026) -
Task-Level Contrastiveness for Cross-Domain Few-Shot Learning
por: Topollai, Kristi, et al.
Publicado: (2025) -
Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
por: Dimlioglu, Tolga, et al.
Publicado: (2025)