The Optimization Landscape of SGD Across the Feature Learning Strength
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Atanasov, Alexander, Meterez, Alexandru, Simon, James B., Pehlevan, Cengiz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
How Feature Learning Can Improve Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
A Dynamical Model of Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
Risk and cross validation in ridge regression with correlated samples
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
Scaling and renormalization in high-dimensional regression
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
Super Consistency of Neural Network Landscapes and Learning Rate Transfer
von: Noci, Lorenzo, et al.
Veröffentlicht: (2024)
von: Noci, Lorenzo, et al.
Veröffentlicht: (2024)
Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge Ensembles
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2023)
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2023)
Transfer Learning in Infinite Width Feature Learning Networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
von: Atanasov, Alexander, et al.
Veröffentlicht: (2025)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2025)
MLPs Learn In-Context on Regression and Classification Tasks
von: Tong, William L., et al.
Veröffentlicht: (2024)
von: Tong, William L., et al.
Veröffentlicht: (2024)
Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
von: Halder, Indranil, et al.
Veröffentlicht: (2025)
von: Halder, Indranil, et al.
Veröffentlicht: (2025)
Learning richness modulates equality reasoning in neural networks
von: Tong, William L., et al.
Veröffentlicht: (2025)
von: Tong, William L., et al.
Veröffentlicht: (2025)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Error Broadcast and Decorrelation as a Potential Artificial and Natural Learning Mechanism
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
Convex Relaxation for Solving Large-Margin Classifiers in Hyperbolic Space
von: Yang, Sheng, et al.
Veröffentlicht: (2024)
von: Yang, Sheng, et al.
Veröffentlicht: (2024)
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
von: Halder, Indranil, et al.
Veröffentlicht: (2026)
von: Halder, Indranil, et al.
Veröffentlicht: (2026)
Universal One-third Time Scaling in Learning Peaked Distributions
von: Liu, Yizhou, et al.
Veröffentlicht: (2026)
von: Liu, Yizhou, et al.
Veröffentlicht: (2026)
An Analytical Theory of Spectral Bias in the Learning Dynamics of Diffusion Models
von: Wang, Binxu, et al.
Veröffentlicht: (2025)
von: Wang, Binxu, et al.
Veröffentlicht: (2025)
Nadaraya-Watson kernel smoothing as a random energy model
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2024)
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2024)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Hyperparameter Transfer with Mixture-of-Expert Layers
von: Jiang, Tianze, et al.
Veröffentlicht: (2026)
von: Jiang, Tianze, et al.
Veröffentlicht: (2026)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2024)
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2024)
Correlative Information Maximization: A Biologically Plausible Approach to Supervised Deep Neural Networks without Weight Symmetry
von: Bozkurt, Bariscan, et al.
Veröffentlicht: (2023)
von: Bozkurt, Bariscan, et al.
Veröffentlicht: (2023)
Infinite Limits of Multi-head Transformer Dynamics
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
A solvable model of learning generative diffusion: theory and insights
von: Cui, Hugo, et al.
Veröffentlicht: (2025)
von: Cui, Hugo, et al.
Veröffentlicht: (2025)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Score Broadcast and Decorrelation: A General Framework for Broadcast-Based Credit Assignment
von: Uzun, Mustafa, et al.
Veröffentlicht: (2026)
von: Uzun, Mustafa, et al.
Veröffentlicht: (2026)
Pixel-Based Similarities as an Alternative to Neural Data for Improving Convolutional Neural Network Adversarial Robustness
von: Attias, Elie, et al.
Veröffentlicht: (2024)
von: Attias, Elie, et al.
Veröffentlicht: (2024)
Boule or Baguette? A Study on Task Topology, Length Generalization, and the Benefit of Reasoning Traces
von: Tong, William L., et al.
Veröffentlicht: (2026)
von: Tong, William L., et al.
Veröffentlicht: (2026)
Pretrain-Test Task Alignment Governs Generalization in In-Context Learning
von: Letey, Mary I., et al.
Veröffentlicht: (2025)
von: Letey, Mary I., et al.
Veröffentlicht: (2025)
Spectral regularization for adversarially-robust representation learning
von: Yang, Sheng, et al.
Veröffentlicht: (2024)
von: Yang, Sheng, et al.
Veröffentlicht: (2024)
Dynamically Learning to Integrate in Recurrent Neural Networks
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
Grokking as the Transition from Lazy to Rich Training Dynamics
von: Kumar, Tanishq, et al.
Veröffentlicht: (2023)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2023)
On the Superlinear Relationship between SGD Noise Covariance and Loss Landscape Curvature
von: Zhang, Yikuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yikuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025) -
How Feature Learning Can Improve Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024) -
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026) -
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
von: Zhao, Rosie, et al.
Veröffentlicht: (2025) -
A Dynamical Model of Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)