From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGD
Fuente:
arXiv
Guardado en:
| Autores principales: | Tsiolis, Konstantinos Christopher, Mousavi-Hosseini, Alireza, Erdogdu, Murat A. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2026)
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2026)
Robust Feature Learning for Multi-Index Models in High Dimensions
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2024)
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2024)
Learning Multi-Index Models with Neural Networks via Mean-Field Langevin Dynamics
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2024)
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2024)
When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2025)
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2025)
A Separation in Heavy-Tailed Sampling: Gaussian vs. Stable Oracles for Proximal Samplers
por: He, Ye, et al.
Publicado: (2024)
por: He, Ye, et al.
Publicado: (2024)
Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
por: Arous, Gérard Ben, et al.
Publicado: (2025)
por: Arous, Gérard Ben, et al.
Publicado: (2025)
Pruning is Optimal for Learning Sparse Features in High-Dimensions
por: Vural, Nuri Mert, et al.
Publicado: (2024)
por: Vural, Nuri Mert, et al.
Publicado: (2024)
Optimal Excess Risk Bounds for Empirical Risk Minimization on $p$-Norm Linear Regression
por: Hanchi, Ayoub El, et al.
Publicado: (2023)
por: Hanchi, Ayoub El, et al.
Publicado: (2023)
On the Efficiency of ERM in Feature Learning
por: Hanchi, Ayoub El, et al.
Publicado: (2024)
por: Hanchi, Ayoub El, et al.
Publicado: (2024)
A Geometric Analysis of PCA
por: Hanchi, Ayoub El, et al.
Publicado: (2025)
por: Hanchi, Ayoub El, et al.
Publicado: (2025)
Mean-Field Langevin Dynamics for Signed Measures via a Bilevel Approach
por: Wang, Guillaume, et al.
Publicado: (2024)
por: Wang, Guillaume, et al.
Publicado: (2024)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
por: Evron, Itay, et al.
Publicado: (2025)
por: Evron, Itay, et al.
Publicado: (2025)
On Fitting Flow Models with Large Sinkhorn Couplings
por: Zhang, Stephen, et al.
Publicado: (2025)
por: Zhang, Stephen, et al.
Publicado: (2025)
Beyond Labeling Oracles: What does it mean to steal ML models?
por: Shafran, Avital, et al.
Publicado: (2023)
por: Shafran, Avital, et al.
Publicado: (2023)
Minimax Linear Regression under the Quantile Risk
por: Hanchi, Ayoub El, et al.
Publicado: (2024)
por: Hanchi, Ayoub El, et al.
Publicado: (2024)
Flow Matching with Semidiscrete Couplings
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2025)
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2025)
From Privacy to Generalization: Linear Max-Information Bounds for DP-SGD
por: Lampert, Christoph H., et al.
Publicado: (2026)
por: Lampert, Christoph H., et al.
Publicado: (2026)
The Marginal Value of Momentum for Small Learning Rate SGD
por: Wang, Runzhe, et al.
Publicado: (2023)
por: Wang, Runzhe, et al.
Publicado: (2023)
Fundamental Limitations of Favorable Privacy-Utility Guarantees for DP-SGD
por: Ertan, Murat Bilgehan, et al.
Publicado: (2026)
por: Ertan, Murat Bilgehan, et al.
Publicado: (2026)
Cumulative Learning Rate Adaptation: Revisiting Path-Based Schedules for SGD and Adam
por: Atamna, Asma, et al.
Publicado: (2025)
por: Atamna, Asma, et al.
Publicado: (2025)
Sampling from the Mean-Field Stationary Distribution
por: Kook, Yunbum, et al.
Publicado: (2024)
por: Kook, Yunbum, et al.
Publicado: (2024)
Analysis of Langevin Monte Carlo from Poincaré to Log-Sobolev
por: Chewi, Sinho, et al.
Publicado: (2021)
por: Chewi, Sinho, et al.
Publicado: (2021)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
por: Khaled, Ahmed, et al.
Publicado: (2025)
por: Khaled, Ahmed, et al.
Publicado: (2025)
Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGD
por: Peng, Ze, et al.
Publicado: (2026)
por: Peng, Ze, et al.
Publicado: (2026)
Generalization and Optimization of SGD with Lookahead
por: Li, Kangcheng, et al.
Publicado: (2025)
por: Li, Kangcheng, et al.
Publicado: (2025)
AutoSGD: Automatic Learning Rate Selection for Stochastic Gradient Descent
por: Surjanovic, Nikola, et al.
Publicado: (2025)
por: Surjanovic, Nikola, et al.
Publicado: (2025)
Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates
por: Glentis, Athanasios, et al.
Publicado: (2026)
por: Glentis, Athanasios, et al.
Publicado: (2026)
Missing-Data-Induced Phase Transitions in Spectral PLS for Multimodal Learning
por: Gjølbye, Anders, et al.
Publicado: (2026)
por: Gjølbye, Anders, et al.
Publicado: (2026)
Trade-off Functions for DP-SGD with Subsampling based on Random Shuffling: Tight Upper and Lower Bounds
por: van Dijk, Marten, et al.
Publicado: (2026)
por: van Dijk, Marten, et al.
Publicado: (2026)
Topology-aware Generalization of Decentralized SGD
por: Zhu, Tongtian, et al.
Publicado: (2022)
por: Zhu, Tongtian, et al.
Publicado: (2022)
Stability and Generalization for Decentralized Markov SGD
por: Wang, Jiahuan, et al.
Publicado: (2026)
por: Wang, Jiahuan, et al.
Publicado: (2026)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
por: Umeda, Hikaru, et al.
Publicado: (2025)
por: Umeda, Hikaru, et al.
Publicado: (2025)
Error Exponent in Agnostic PAC Learning
por: Hendel, Adi, et al.
Publicado: (2024)
por: Hendel, Adi, et al.
Publicado: (2024)
Online Learning and Information Exponents: On The Importance of Batch size, and Time/Complexity Tradeoffs
por: Arnaboldi, Luca, et al.
Publicado: (2024)
por: Arnaboldi, Luca, et al.
Publicado: (2024)
Effect of Random Learning Rate: Theoretical Analysis of SGD Dynamics in Non-Convex Optimization via Stationary Distribution
por: Yoshida, Naoki, et al.
Publicado: (2024)
por: Yoshida, Naoki, et al.
Publicado: (2024)
Phases of Muon: When Muon Eclipses SignSGD
por: Paquette, Elliot, et al.
Publicado: (2026)
por: Paquette, Elliot, et al.
Publicado: (2026)
Unveiling High-Probability Generalization in Decentralized SGD
por: Wang, Jiahuan, et al.
Publicado: (2026)
por: Wang, Jiahuan, et al.
Publicado: (2026)
Learning Orthogonal Multi-Index Models: A Fine-Grained Information Exponent Analysis
por: Ren, Yunwei, et al.
Publicado: (2024)
por: Ren, Yunwei, et al.
Publicado: (2024)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
por: Kovalev, Dmitry, et al.
Publicado: (2025)
por: Kovalev, Dmitry, et al.
Publicado: (2025)
Shuffling the Data, Stretching the Step-size: Sharper Bias in constant step-size SGD
por: Emmanouilidis, Konstantinos, et al.
Publicado: (2026)
por: Emmanouilidis, Konstantinos, et al.
Publicado: (2026)
Ejemplares similares
-
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2026) -
Robust Feature Learning for Multi-Index Models in High Dimensions
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2024) -
Learning Multi-Index Models with Neural Networks via Mean-Field Langevin Dynamics
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2024) -
When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2025) -
A Separation in Heavy-Tailed Sampling: Gaussian vs. Stable Oracles for Proximal Samplers
por: He, Ye, et al.
Publicado: (2024)