Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training
Fuente:
arXiv
Saved in:
| Main Authors: | Naganuma, Hiroki, Zhang, Xinzhi, Yue, Man-Chung, Mitliagkas, Ioannis, Witte, Philipp A., Hewett, Russell J., Lee, Yin Tat |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
by: Naganuma, Hiroki, et al.
Published: (2023)
by: Naganuma, Hiroki, et al.
Published: (2023)
Asynchronous Local-SGD Training for Language Modeling
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
Revisiting Generalization Measures Beyond IID: An Empirical Study under Distributional Shift
by: Nakai, Sora, et al.
Published: (2026)
by: Nakai, Sora, et al.
Published: (2026)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
Orth-Dion: Eliminating Geometric Mismatch in Distributed Low-Rank Spectral Optimization
by: Nakamori, Tatsuhiro, et al.
Published: (2026)
by: Nakamori, Tatsuhiro, et al.
Published: (2026)
Birch SGD: A Tree Graph Framework for Local and Asynchronous SGD Methods
by: Tyurin, Alexander, et al.
Published: (2025)
by: Tyurin, Alexander, et al.
Published: (2025)
Dual-Delayed Asynchronous SGD for Arbitrarily Heterogeneous Data
by: Wang, Xiaolu, et al.
Published: (2024)
by: Wang, Xiaolu, et al.
Published: (2024)
Generating Tabular Data Using Heterogeneous Sequential Feature Forest Flow Matching
by: Akazan, Ange-Clément, et al.
Published: (2024)
by: Akazan, Ange-Clément, et al.
Published: (2024)
HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model Training
by: Kim, Geon-Woo, et al.
Published: (2025)
by: Kim, Geon-Woo, et al.
Published: (2025)
Ordered Momentum for Asynchronous SGD
by: Shi, Chang-Wei, et al.
Published: (2024)
by: Shi, Chang-Wei, et al.
Published: (2024)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Performative Prediction with Neural Networks
by: Mofakhami, Mehrnaz, et al.
Published: (2023)
by: Mofakhami, Mehrnaz, et al.
Published: (2023)
First Provably Optimal Asynchronous SGD for Homogeneous and Heterogeneous Data
by: Maranjyan, Artavazd
Published: (2026)
by: Maranjyan, Artavazd
Published: (2026)
Biased Local SGD for Efficient Deep Learning on Heterogeneous Systems
by: Lim, Jihyun, et al.
Published: (2025)
by: Lim, Jihyun, et al.
Published: (2025)
Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
by: Inane, Ahmed Mehdi, et al.
Published: (2026)
by: Inane, Ahmed Mehdi, et al.
Published: (2026)
Bringing Order to Asynchronous SGD: Towards Optimality under Data-Dependent Delays with Momentum
by: Dahan, Tehila, et al.
Published: (2026)
by: Dahan, Tehila, et al.
Published: (2026)
A Derandomization Framework for Structure Discovery: Applications in Neural Networks and Beyond
by: Tsikouras, Nikos, et al.
Published: (2025)
by: Tsikouras, Nikos, et al.
Published: (2025)
Towards Understanding Variants of Invariant Risk Minimization through the Lens of Calibration
by: Yoshida, Kotaro, et al.
Published: (2024)
by: Yoshida, Kotaro, et al.
Published: (2024)
MindFlayer SGD: Efficient Parallel SGD in the Presence of Heterogeneous and Random Worker Compute Times
by: Maranjyan, Artavazd, et al.
Published: (2024)
by: Maranjyan, Artavazd, et al.
Published: (2024)
Towards efficient representation identification in supervised learning
by: Ahuja, Kartik, et al.
Published: (2022)
by: Ahuja, Kartik, et al.
Published: (2022)
Geometric Insights into Focal Loss: Reducing Curvature for Enhanced Model Calibration
by: Kimura, Masanari, et al.
Published: (2024)
by: Kimura, Masanari, et al.
Published: (2024)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
by: Tyurin, Alexander, et al.
Published: (2024)
by: Tyurin, Alexander, et al.
Published: (2024)
Rescaled Asynchronous SGD: Optimal Distributed Optimization under Data and System Heterogeneity
by: Mahran, Ammar, et al.
Published: (2026)
by: Mahran, Ammar, et al.
Published: (2026)
Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation
by: Mahajan, Divyat, et al.
Published: (2022)
by: Mahajan, Divyat, et al.
Published: (2022)
Ringleader ASGD: The First Asynchronous SGD with Optimal Time Complexity under Data Heterogeneity
by: Maranjyan, Artavazd, et al.
Published: (2025)
by: Maranjyan, Artavazd, et al.
Published: (2025)
A Study of Condition Numbers for First-Order Optimization
by: Guille-Escuret, Charles, et al.
Published: (2020)
by: Guille-Escuret, Charles, et al.
Published: (2020)
Data Debugging is NP-hard for Classifiers Trained with SGD
by: Guo, Zizheng, et al.
Published: (2024)
by: Guo, Zizheng, et al.
Published: (2024)
Ringmaster ASGD: The First Asynchronous SGD with Optimal Time Complexity
by: Maranjyan, Artavazd, et al.
Published: (2025)
by: Maranjyan, Artavazd, et al.
Published: (2025)
Exploring Scaling Laws for Local SGD in Large Language Model Training
by: He, Qiaozhi, et al.
Published: (2024)
by: He, Qiaozhi, et al.
Published: (2024)
DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
by: Yoshida, Kotaro, et al.
Published: (2025)
by: Yoshida, Kotaro, et al.
Published: (2025)
Enhancing Generative Networks for Chest Anomaly Localization through Automatic Registration-Based Unpaired-to-Pseudo-Paired Training Data Translation
by: Kim, Kyungsu, et al.
Published: (2022)
by: Kim, Kyungsu, et al.
Published: (2022)
Convergence Bound and Critical Batch Size of Muon Optimizer
by: Sato, Naoki, et al.
Published: (2025)
by: Sato, Naoki, et al.
Published: (2025)
Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates
by: Glentis, Athanasios, et al.
Published: (2026)
by: Glentis, Athanasios, et al.
Published: (2026)
Performance Control in Early Exiting to Deploy Large Models at the Same Cost of Smaller Ones
by: Mofakhami, Mehrnaz, et al.
Published: (2024)
by: Mofakhami, Mehrnaz, et al.
Published: (2024)
Perfect Parallelization in Mini-Batch SGD with Classical Momentum Acceleration
by: Garg, Sachin, et al.
Published: (2026)
by: Garg, Sachin, et al.
Published: (2026)
Generalizing to unseen domains via distribution matching
by: Albuquerque, Isabela, et al.
Published: (2019)
by: Albuquerque, Isabela, et al.
Published: (2019)
SGD Jittering: A Training Strategy for Robust and Accurate Model-Based Architectures
by: Guan, Peimeng, et al.
Published: (2024)
by: Guan, Peimeng, et al.
Published: (2024)
AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
by: Ajanthan, Thalaiyasingam, et al.
Published: (2026)
by: Ajanthan, Thalaiyasingam, et al.
Published: (2026)
Upper and Lower Bounds on the Smoothed Complexity of the Simplex Method
by: Huiberts, Sophie, et al.
Published: (2022)
by: Huiberts, Sophie, et al.
Published: (2022)
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
by: Chen, Ling, et al.
Published: (2026)
by: Chen, Ling, et al.
Published: (2026)
Similar Items
-
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
by: Naganuma, Hiroki, et al.
Published: (2023) -
Asynchronous Local-SGD Training for Language Modeling
by: Liu, Bo, et al.
Published: (2024) -
Revisiting Generalization Measures Beyond IID: An Empirical Study under Distributional Shift
by: Nakai, Sora, et al.
Published: (2026) -
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
by: Naganuma, Hiroki, et al.
Published: (2026) -
Orth-Dion: Eliminating Geometric Mismatch in Distributed Low-Rank Spectral Optimization
by: Nakamori, Tatsuhiro, et al.
Published: (2026)