From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Ziyan, Zhou, Ding-Xuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improved Scaling Laws in Linear Regression via Data Reuse
por: Lin, Licong, et al.
Publicado: (2025)
por: Lin, Licong, et al.
Publicado: (2025)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
por: Kim, Jihwan, et al.
Publicado: (2026)
por: Kim, Jihwan, et al.
Publicado: (2026)
Accelerating Single-Pass SGD for Generalized Linear Prediction
por: Chen, Qian, et al.
Publicado: (2026)
por: Chen, Qian, et al.
Publicado: (2026)
Perfect Parallelization in Mini-Batch SGD with Classical Momentum Acceleration
por: Garg, Sachin, et al.
Publicado: (2026)
por: Garg, Sachin, et al.
Publicado: (2026)
Data Deletion for Linear Regression with Noisy SGD
por: Xia, Zhangjie, et al.
Publicado: (2024)
por: Xia, Zhangjie, et al.
Publicado: (2024)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
por: Kovačević, Filip, et al.
Publicado: (2026)
por: Kovačević, Filip, et al.
Publicado: (2026)
Heavy-Tailed Linear Bandits: Huber Regression with One-Pass Update
por: Wang, Jing, et al.
Publicado: (2025)
por: Wang, Jing, et al.
Publicado: (2025)
Scaling Laws for Precision in High-Dimensional Linear Regression
por: Zhang, Dechen, et al.
Publicado: (2026)
por: Zhang, Dechen, et al.
Publicado: (2026)
Private Sketches for Linear Regression
por: Das, Shrutimoy, et al.
Publicado: (2025)
por: Das, Shrutimoy, et al.
Publicado: (2025)
Scaling Law for Stochastic Gradient Descent in Quadratically Parameterized Linear Regression
por: Ding, Shihong, et al.
Publicado: (2025)
por: Ding, Shihong, et al.
Publicado: (2025)
Scaling Laws in Linear Regression: Compute, Parameters, and Data
por: Lin, Licong, et al.
Publicado: (2024)
por: Lin, Licong, et al.
Publicado: (2024)
Hidden State Differential Private Mini-Batch Block Coordinate Descent for Multi-convexity Optimization
por: Chen, Ding, et al.
Publicado: (2024)
por: Chen, Ding, et al.
Publicado: (2024)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
por: Meterez, Alexandru, et al.
Publicado: (2025)
por: Meterez, Alexandru, et al.
Publicado: (2025)
Enhancing SignSGD: Small-Batch Convergence Analysis and a Hybrid Switching Strategy
por: Chen, Haoran, et al.
Publicado: (2026)
por: Chen, Haoran, et al.
Publicado: (2026)
Online Covariance Estimation in Averaged SGD: Improved Batch-Mean Rates and Minimax Optimality via Trajectory Regression
por: Ni, Yijin, et al.
Publicado: (2026)
por: Ni, Yijin, et al.
Publicado: (2026)
Scaling Federated Linear Contextual Bandits via Sketching
por: Yang, Hantao, et al.
Publicado: (2026)
por: Yang, Hantao, et al.
Publicado: (2026)
Statistical Inference for Linear Functionals of Online SGD in High-dimensional Linear Regression
por: Agrawalla, Bhavya, et al.
Publicado: (2023)
por: Agrawalla, Bhavya, et al.
Publicado: (2023)
Understanding SGD with Exponential Moving Average: A Case Study in Linear Regression
por: Li, Xuheng, et al.
Publicado: (2025)
por: Li, Xuheng, et al.
Publicado: (2025)
On the (Generative) Linear Sketching Problem
por: Yuan, Xinyu, et al.
Publicado: (2026)
por: Yuan, Xinyu, et al.
Publicado: (2026)
Exploring Scaling Laws for Local SGD in Large Language Model Training
por: He, Qiaozhi, et al.
Publicado: (2024)
por: He, Qiaozhi, et al.
Publicado: (2024)
Bayesian Data Sketching for Varying Coefficient Regression Models
por: Guhaniyogi, Rajarshi, et al.
Publicado: (2025)
por: Guhaniyogi, Rajarshi, et al.
Publicado: (2025)
Debiasing Mini-Batch Quadratics for Applications in Deep Learning
por: Tatzel, Lukas, et al.
Publicado: (2024)
por: Tatzel, Lukas, et al.
Publicado: (2024)
From Privacy to Generalization: Linear Max-Information Bounds for DP-SGD
por: Lampert, Christoph H., et al.
Publicado: (2026)
por: Lampert, Christoph H., et al.
Publicado: (2026)
Towards Scaling Laws for Symbolic Regression
por: Otte, David, et al.
Publicado: (2025)
por: Otte, David, et al.
Publicado: (2025)
Generalized Linear Bandits: Almost Optimal Regret with One-Pass Update
por: Zhang, Yu-Jie, et al.
Publicado: (2025)
por: Zhang, Yu-Jie, et al.
Publicado: (2025)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
por: Evron, Itay, et al.
Publicado: (2025)
por: Evron, Itay, et al.
Publicado: (2025)
Scaling Law for Language Models Training Considering Batch Size
por: Shuai, Xian, et al.
Publicado: (2024)
por: Shuai, Xian, et al.
Publicado: (2024)
Mini-Batch Kernel $k$-means
por: Jourdan, Ben, et al.
Publicado: (2024)
por: Jourdan, Ben, et al.
Publicado: (2024)
From Message-Passing to Linearized Graph Sequence Models
por: Mathys, Joël, et al.
Publicado: (2026)
por: Mathys, Joël, et al.
Publicado: (2026)
Mini-Batch Class Composition Bias in Link Prediction
por: Maguire, Kieran, et al.
Publicado: (2026)
por: Maguire, Kieran, et al.
Publicado: (2026)
Batch List-Decodable Linear Regression via Higher Moments
por: Diakonikolas, Ilias, et al.
Publicado: (2025)
por: Diakonikolas, Ilias, et al.
Publicado: (2025)
Can Microcanonical Langevin Dynamics Leverage Mini-Batch Gradient Noise?
por: Sommer, Emanuel, et al.
Publicado: (2026)
por: Sommer, Emanuel, et al.
Publicado: (2026)
SGD for Variational Inference: Tackling Unbounded Variance via Preconditioning and Dynamic Batching
por: Labarrière, Hippolyte, et al.
Publicado: (2026)
por: Labarrière, Hippolyte, et al.
Publicado: (2026)
Probability Passing for Graph Neural Networks: Graph Structure and Representations Joint Learning
por: Wang, Ziyan, et al.
Publicado: (2024)
por: Wang, Ziyan, et al.
Publicado: (2024)
Exact Mean Square Linear Stability Analysis for SGD
por: Mulayoff, Rotem, et al.
Publicado: (2023)
por: Mulayoff, Rotem, et al.
Publicado: (2023)
Neural Scaling Laws for Deep Regression
por: Cadez, Tilen, et al.
Publicado: (2025)
por: Cadez, Tilen, et al.
Publicado: (2025)
Full-Graph vs. Mini-Batch Training: Comprehensive Analysis from a Batch Size and Fan-Out Size Perspective
por: Liu, Mengfan, et al.
Publicado: (2026)
por: Liu, Mengfan, et al.
Publicado: (2026)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
por: Umeda, Hikaru, et al.
Publicado: (2025)
por: Umeda, Hikaru, et al.
Publicado: (2025)
Hierarchical Rectified Flow Matching with Mini-Batch Couplings
por: Zhang, Yichi, et al.
Publicado: (2025)
por: Zhang, Yichi, et al.
Publicado: (2025)
Deep Sketched Output Kernel Regression for Structured Prediction
por: Ahmad, Tamim El, et al.
Publicado: (2024)
por: Ahmad, Tamim El, et al.
Publicado: (2024)
Ejemplares similares
-
Improved Scaling Laws in Linear Regression via Data Reuse
por: Lin, Licong, et al.
Publicado: (2025) -
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
por: Kim, Jihwan, et al.
Publicado: (2026) -
Accelerating Single-Pass SGD for Generalized Linear Prediction
por: Chen, Qian, et al.
Publicado: (2026) -
Perfect Parallelization in Mini-Batch SGD with Classical Momentum Acceleration
por: Garg, Sachin, et al.
Publicado: (2026) -
Data Deletion for Linear Regression with Noisy SGD
por: Xia, Zhangjie, et al.
Publicado: (2024)