A Mechanism Study of Delayed Loss Spikes in Batch-Normalized Linear Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Peifeng, Fang, Wenyi, Zheng, Yang, Zou, Difan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Online Linear Programming with Batching
by: Xu, Haoran, et al.
Published: (2024)
by: Xu, Haoran, et al.
Published: (2024)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
Faster Stochastic Optimization with Arbitrary Delays via Asynchronous Mini-Batching
by: Attia, Amit, et al.
Published: (2024)
by: Attia, Amit, et al.
Published: (2024)
Faster Sampling without Isoperimetry via Diffusion-based Monte Carlo
by: Huang, Xunpeng, et al.
Published: (2024)
by: Huang, Xunpeng, et al.
Published: (2024)
An Improved Analysis of Langevin Algorithms with Prior Diffusion for Non-Log-Concave Sampling
by: Huang, Xunpeng, et al.
Published: (2024)
by: Huang, Xunpeng, et al.
Published: (2024)
Adaptive Gradient Normalization and Independent Sampling for (Stochastic) Generalized-Smooth Optimization
by: Yang, Yufeng, et al.
Published: (2024)
by: Yang, Yufeng, et al.
Published: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Smart Surrogate Losses for Contextual Stochastic Linear Optimization with Robust Constraints
by: Im, Hyungki, et al.
Published: (2025)
by: Im, Hyungki, et al.
Published: (2025)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
by: Hanqing, Liu, et al.
Published: (2026)
by: Hanqing, Liu, et al.
Published: (2026)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Optimistic Online-to-Batch Conversions for Accelerated Convergence and Universality
by: Yan, Yu-Hu, et al.
Published: (2025)
by: Yan, Yu-Hu, et al.
Published: (2025)
Bilevel Models for Adversarial Learning and A Case Study
by: Zheng, Yutong, et al.
Published: (2025)
by: Zheng, Yutong, et al.
Published: (2025)
Batched First-Order Methods for Parallel LP Solving in MIP
by: Blin, Nicolas, et al.
Published: (2026)
by: Blin, Nicolas, et al.
Published: (2026)
Accelerating Single-Pass SGD for Generalized Linear Prediction
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
The Implicit Bias of Heterogeneity towards Invariance: A Study of Multi-Environment Matrix Sensing
by: Xu, Yang, et al.
Published: (2024)
by: Xu, Yang, et al.
Published: (2024)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
by: Oowada, Kanata, et al.
Published: (2025)
by: Oowada, Kanata, et al.
Published: (2025)
Nearly Optimal Linear Convergence of Stochastic Primal-Dual Methods for Linear Programming
by: Lu, Haihao, et al.
Published: (2021)
by: Lu, Haihao, et al.
Published: (2021)
Predictor-Based Output-Feedback Control of Linear Systems with Time-Varying Input and Measurement Delays via Neural-Approximated Prediction Horizons
by: Bhan, Luke, et al.
Published: (2026)
by: Bhan, Luke, et al.
Published: (2026)
Quantitative Convergence Analysis of Projected Stochastic Gradient Descent for Non-Convex Losses via the Goldstein Subdifferential
by: Zheng, Yuping, et al.
Published: (2025)
by: Zheng, Yuping, et al.
Published: (2025)
SGD for Variational Inference: Tackling Unbounded Variance via Preconditioning and Dynamic Batching
by: Labarrière, Hippolyte, et al.
Published: (2026)
by: Labarrière, Hippolyte, et al.
Published: (2026)
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024)
by: Umeda, Hikaru, et al.
Published: (2024)
Stochastic Variance-Reduced Newton: Accelerating Finite-Sum Minimization with Large Batches
by: Dereziński, Michał
Published: (2022)
by: Dereziński, Michał
Published: (2022)
Regret Bounds for Episodic Risk-Sensitive Linear Quadratic Regulator
by: Xu, Wenhao, et al.
Published: (2024)
by: Xu, Wenhao, et al.
Published: (2024)
Near-optimal and Efficient First-Order Algorithm for Multi-Task Learning with Shared Linear Representation
by: Ding, Shihong, et al.
Published: (2026)
by: Ding, Shihong, et al.
Published: (2026)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Understanding SGD with Exponential Moving Average: A Case Study in Linear Regression
by: Li, Xuheng, et al.
Published: (2025)
by: Li, Xuheng, et al.
Published: (2025)
Fair Generalized Linear Mixed Models
by: Burgard, Jan Pablo, et al.
Published: (2024)
by: Burgard, Jan Pablo, et al.
Published: (2024)
A Semantic-Loss Function Modeling Framework With Task-Oriented Machine Learning Perspectives
by: Nguyen, Ti Ti, et al.
Published: (2025)
by: Nguyen, Ti Ti, et al.
Published: (2025)
From Sequential Nodes to GPU Batches: Parallel Branch and Bound for Optimal $k$-Sparse GLMs
by: Liu, Jiachang, et al.
Published: (2026)
by: Liu, Jiachang, et al.
Published: (2026)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
by: Harada, Hinata, et al.
Published: (2024)
by: Harada, Hinata, et al.
Published: (2024)
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
by: Imaizumi, Kento, et al.
Published: (2025)
by: Imaizumi, Kento, et al.
Published: (2025)
Modified Meta-Thompson Sampling for Linear Bandits and Its Bayes Regret Analysis
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
by: Wang, Jinbo, et al.
Published: (2026)
by: Wang, Jinbo, et al.
Published: (2026)
Asynchronous Distributed Optimization with Delay-free Parameters
by: Wu, Xuyang, et al.
Published: (2023)
by: Wu, Xuyang, et al.
Published: (2023)
Linear Model Extraction via Factual and Counterfactual Queries
by: Otto, Daan, et al.
Published: (2026)
by: Otto, Daan, et al.
Published: (2026)
Scalable Approximate Algorithms for Optimal Transport Linear Models
by: Kacprzak, Tomasz, et al.
Published: (2025)
by: Kacprzak, Tomasz, et al.
Published: (2025)
Adaptive Delayed-Update Cyclic Algorithm for Variational Inequalities
by: Wei, Yi, et al.
Published: (2026)
by: Wei, Yi, et al.
Published: (2026)
Similar Items
-
Online Linear Programming with Batching
by: Xu, Haoran, et al.
Published: (2024) -
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024) -
Faster Stochastic Optimization with Arbitrary Delays via Asynchronous Mini-Batching
by: Attia, Amit, et al.
Published: (2024) -
Faster Sampling without Isoperimetry via Diffusion-based Monte Carlo
by: Huang, Xunpeng, et al.
Published: (2024) -
An Improved Analysis of Langevin Algorithms with Prior Diffusion for Non-Log-Concave Sampling
by: Huang, Xunpeng, et al.
Published: (2024)