A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Kühn, Marcel, Thelge, Yoon, Rosenow, Bernd |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Anti-Correlated Noise in Epoch-Based Stochastic Gradient Descent: Implications for Weight Variances in Flat Directions
by: Kühn, Marcel, et al.
Published: (2023)
by: Kühn, Marcel, et al.
Published: (2023)
Boundary between noise and information applied to filtering neural network weight matrices
by: Staats, Max, et al.
Published: (2022)
by: Staats, Max, et al.
Published: (2022)
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
by: Staats, Max, et al.
Published: (2024)
by: Staats, Max, et al.
Published: (2024)
Enhancing Noise-Robust Losses for Large-Scale Noisy Data Learning
by: Staats, Max, et al.
Published: (2023)
by: Staats, Max, et al.
Published: (2023)
Continuous Specialization Transition in the Soft Committee Machine with ReLU Activation
by: Afanah, Assem, et al.
Published: (2026)
by: Afanah, Assem, et al.
Published: (2026)
Unified Description of Learning Dynamics in the Soft Committee Machine from Finite to Ultra-Wide Regimes
by: Afanah, Assem, et al.
Published: (2025)
by: Afanah, Assem, et al.
Published: (2025)
Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
by: Duranthon, O., et al.
Published: (2025)
by: Duranthon, O., et al.
Published: (2025)
BBP Phase Transition for an Extensive Number of Outliers
by: Forner, Niklas, et al.
Published: (2025)
by: Forner, Niklas, et al.
Published: (2025)
A Federated Many-to-One Hopfield model for associative Neural Networks
by: Alessandrelli, Andrea, et al.
Published: (2026)
by: Alessandrelli, Andrea, et al.
Published: (2026)
A Dynamical Model of Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks
by: Montanari, Andrea, et al.
Published: (2025)
by: Montanari, Andrea, et al.
Published: (2025)
Grokking as a First Order Phase Transition in Two Layer Networks
by: Rubin, Noa, et al.
Published: (2023)
by: Rubin, Noa, et al.
Published: (2023)
The Quantization Model of Neural Scaling
by: Michaud, Eric J., et al.
Published: (2023)
by: Michaud, Eric J., et al.
Published: (2023)
Explaining Neural Scaling Laws
by: Bahri, Yasaman, et al.
Published: (2021)
by: Bahri, Yasaman, et al.
Published: (2021)
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
by: Arnaboldi, Luca, et al.
Published: (2025)
by: Arnaboldi, Luca, et al.
Published: (2025)
Scaling and renormalization in high-dimensional regression
by: Atanasov, Alexander, et al.
Published: (2024)
by: Atanasov, Alexander, et al.
Published: (2024)
Graph Neural Network Approach to Predicting Magnetization in Quasi-One-Dimensional Ising Systems
by: Slavin, V., et al.
Published: (2025)
by: Slavin, V., et al.
Published: (2025)
From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
by: Rubin, Noa, et al.
Published: (2025)
by: Rubin, Noa, et al.
Published: (2025)
Neural Scaling Laws Rooted in the Data Distribution
by: Brill, Ari
Published: (2024)
by: Brill, Ari
Published: (2024)
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
by: Kalra, Dayal Singh, et al.
Published: (2024)
by: Kalra, Dayal Singh, et al.
Published: (2024)
How Feature Learning Can Improve Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Statistical Mechanics Calculations Using Variational Autoregressive Networks and Quantum Annealing
by: Tamura, Yuta, et al.
Published: (2024)
by: Tamura, Yuta, et al.
Published: (2024)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
by: Ruben, Benjamin S., et al.
Published: (2024)
by: Ruben, Benjamin S., et al.
Published: (2024)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Dropout Universality: Scaling Laws and Optimal Scheduling at the Edge-of-Chaos
by: Sarmiento, Lucas Fernandez
Published: (2026)
by: Sarmiento, Lucas Fernandez
Published: (2026)
A Generative Diffusion Model for Amorphous Materials
by: Yang, Kai, et al.
Published: (2025)
by: Yang, Kai, et al.
Published: (2025)
A unified theory of feature learning in RNNs and DNNs
by: Bauer, Jan P., et al.
Published: (2026)
by: Bauer, Jan P., et al.
Published: (2026)
A Theory of Saddle Escape in Deep Nonlinear Networks
by: Rawal, Divit, et al.
Published: (2026)
by: Rawal, Divit, et al.
Published: (2026)
A universal approximation theorem for nonlinear resistive networks
by: Scellier, Benjamin, et al.
Published: (2023)
by: Scellier, Benjamin, et al.
Published: (2023)
A solvable model of learning generative diffusion: theory and insights
by: Cui, Hugo, et al.
Published: (2025)
by: Cui, Hugo, et al.
Published: (2025)
DCEM: A deep complementary energy method for solid mechanics
by: Wang, Yizheng, et al.
Published: (2023)
by: Wang, Yizheng, et al.
Published: (2023)
A Random-Matrix Criterion for Initializing Gated Recurrent Neural Networks
by: Fioratti, Tommaso, et al.
Published: (2026)
by: Fioratti, Tommaso, et al.
Published: (2026)
Towards Understanding Inductive Bias in Transformers: A View From Infinity
by: Lavie, Itay, et al.
Published: (2024)
by: Lavie, Itay, et al.
Published: (2024)
A High Dimensional Statistical Model for Adversarial Training: Geometry and Trade-Offs
by: Tanner, Kasimir, et al.
Published: (2024)
by: Tanner, Kasimir, et al.
Published: (2024)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
by: Dandi, Yatin, et al.
Published: (2026)
by: Dandi, Yatin, et al.
Published: (2026)
Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
by: Dawid, Anna, et al.
Published: (2023)
by: Dawid, Anna, et al.
Published: (2023)
Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens
by: Erba, Vittorio, et al.
Published: (2024)
by: Erba, Vittorio, et al.
Published: (2024)
Precise Dynamics of Diagonal Linear Networks: A Unifying Analysis by Dynamical Mean-Field Theory
by: Nishiyama, Sota, et al.
Published: (2025)
by: Nishiyama, Sota, et al.
Published: (2025)
Similar Items
-
Anti-Correlated Noise in Epoch-Based Stochastic Gradient Descent: Implications for Weight Variances in Flat Directions
by: Kühn, Marcel, et al.
Published: (2023) -
Boundary between noise and information applied to filtering neural network weight matrices
by: Staats, Max, et al.
Published: (2022) -
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
by: Staats, Max, et al.
Published: (2024) -
Enhancing Noise-Robust Losses for Large-Scale Noisy Data Learning
by: Staats, Max, et al.
Published: (2023) -
Continuous Specialization Transition in the Soft Committee Machine with ReLU Activation
by: Afanah, Assem, et al.
Published: (2026)