Online Learning and Information Exponents: On The Importance of Batch size, and Time/Complexity Tradeoffs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Arnaboldi, Luca, Dandi, Yatin, Krzakala, Florent, Loureiro, Bruno, Pesce, Luca, Stephan, Ludovic |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions
von: Arnaboldi, Luca, et al.
Veröffentlicht: (2024)
von: Arnaboldi, Luca, et al.
Veröffentlicht: (2024)
The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents
von: Dandi, Yatin, et al.
Veröffentlicht: (2024)
von: Dandi, Yatin, et al.
Veröffentlicht: (2024)
How Two-Layer Neural Networks Learn, One (Giant) Step at a Time
von: Dandi, Yatin, et al.
Veröffentlicht: (2023)
von: Dandi, Yatin, et al.
Veröffentlicht: (2023)
Deep Learning of Compositional Targets with Hierarchical Spectral Methods
von: Tabanelli, Hugo, et al.
Veröffentlicht: (2026)
von: Tabanelli, Hugo, et al.
Veröffentlicht: (2026)
The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
von: Dandi, Yatin, et al.
Veröffentlicht: (2025)
von: Dandi, Yatin, et al.
Veröffentlicht: (2025)
Escaping mediocrity: how two-layer networks learn hard generalized linear models with SGD
von: Arnaboldi, Luca, et al.
Veröffentlicht: (2023)
von: Arnaboldi, Luca, et al.
Veröffentlicht: (2023)
Universality laws for Gaussian mixtures in generalized linear models
von: Dandi, Yatin, et al.
Veröffentlicht: (2023)
von: Dandi, Yatin, et al.
Veröffentlicht: (2023)
A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities
von: Dandi, Yatin, et al.
Veröffentlicht: (2024)
von: Dandi, Yatin, et al.
Veröffentlicht: (2024)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
von: Dandi, Yatin, et al.
Veröffentlicht: (2026)
von: Dandi, Yatin, et al.
Veröffentlicht: (2026)
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
von: Arnaboldi, Luca, et al.
Veröffentlicht: (2025)
von: Arnaboldi, Luca, et al.
Veröffentlicht: (2025)
Asymptotics of feature learning in two-layer networks after one gradient-step
von: Cui, Hugo, et al.
Veröffentlicht: (2024)
von: Cui, Hugo, et al.
Veröffentlicht: (2024)
Optimal Spectral Transitions in High-Dimensional Multi-Index Models
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2025)
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2025)
Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model
von: Wortsman-Zurich, Arie, et al.
Veröffentlicht: (2026)
von: Wortsman-Zurich, Arie, et al.
Veröffentlicht: (2026)
Provable Learning of Random Hierarchy Models and Hierarchical Shallow-to-Deep Chaining
von: Ren, Yunwei, et al.
Veröffentlicht: (2026)
von: Ren, Yunwei, et al.
Veröffentlicht: (2026)
Fundamental computational limits of weak learnability in high-dimensional multi-index models
von: Troiani, Emanuele, et al.
Veröffentlicht: (2024)
von: Troiani, Emanuele, et al.
Veröffentlicht: (2024)
Sampling with flows, diffusion and autoregressive neural networks: A spin-glass perspective
von: Ghio, Davide, et al.
Veröffentlicht: (2023)
von: Ghio, Davide, et al.
Veröffentlicht: (2023)
A Noise Sensitivity Exponent Controls Large Statistical-to-Computational Gaps in Single- and Multi-Index Models
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2026)
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2026)
Asymptotics of Non-Convex Generalized Linear Models in High-Dimensions: A proof of the replica formula
von: Vilucchio, Matteo, et al.
Veröffentlicht: (2025)
von: Vilucchio, Matteo, et al.
Veröffentlicht: (2025)
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
von: Troiani, Emanuele, et al.
Veröffentlicht: (2025)
von: Troiani, Emanuele, et al.
Veröffentlicht: (2025)
Gaussian Universality of Perceptrons with Random Labels
von: Gerace, Federica, et al.
Veröffentlicht: (2022)
von: Gerace, Federica, et al.
Veröffentlicht: (2022)
Optimal scaling laws in learning hierarchical multi-index models
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2026)
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2026)
A Gentle Introduction to Gradient-Based Optimization and Variational Inequalities for Machine Learning
von: Wadia, Neha S., et al.
Veröffentlicht: (2023)
von: Wadia, Neha S., et al.
Veröffentlicht: (2023)
A High Dimensional Statistical Model for Adversarial Training: Geometry and Trade-Offs
von: Tanner, Kasimir, et al.
Veröffentlicht: (2024)
von: Tanner, Kasimir, et al.
Veröffentlicht: (2024)
Rigorous Asymptotics for First-Order Algorithms Through the Dynamical Cavity Method
von: Dandi, Yatin, et al.
Veröffentlicht: (2026)
von: Dandi, Yatin, et al.
Veröffentlicht: (2026)
Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression
von: Clarté, Lucas, et al.
Veröffentlicht: (2024)
von: Clarté, Lucas, et al.
Veröffentlicht: (2024)
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
von: Chaffin, Antoine, et al.
Veröffentlicht: (2026)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2026)
Regret-Oracle Complexity Tradeoffs in Agnostic Online Learning
von: Attias, Idan, et al.
Veröffentlicht: (2026)
von: Attias, Idan, et al.
Veröffentlicht: (2026)
Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions
von: Xu, Yizhou, et al.
Veröffentlicht: (2025)
von: Xu, Yizhou, et al.
Veröffentlicht: (2025)
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2025)
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2025)
Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers
von: Kezins, Nikita, et al.
Veröffentlicht: (2026)
von: Kezins, Nikita, et al.
Veröffentlicht: (2026)
Fundamental limits of Non-Linear Low-Rank Matrix Estimation
von: Mergny, Pierre, et al.
Veröffentlicht: (2024)
von: Mergny, Pierre, et al.
Veröffentlicht: (2024)
A phase transition between positional and semantic learning in a solvable model of dot-product attention
von: Cui, Hugo, et al.
Veröffentlicht: (2024)
von: Cui, Hugo, et al.
Veröffentlicht: (2024)
Asymptotic Characterisation of Robust Empirical Risk Minimisation Performance in the Presence of Outliers
von: Vilucchio, Matteo, et al.
Veröffentlicht: (2023)
von: Vilucchio, Matteo, et al.
Veröffentlicht: (2023)
Computational Thresholds in Multi-Modal Learning via the Spiked Matrix-Tensor Model
von: Tabanelli, Hugo, et al.
Veröffentlicht: (2025)
von: Tabanelli, Hugo, et al.
Veröffentlicht: (2025)
Analysis of learning a flow-based generative model from limited sample complexity
von: Cui, Hugo, et al.
Veröffentlicht: (2023)
von: Cui, Hugo, et al.
Veröffentlicht: (2023)
Spectral Phase Transition and Optimal PCA in Block-Structured Spiked models
von: Mergny, Pierre, et al.
Veröffentlicht: (2024)
von: Mergny, Pierre, et al.
Veröffentlicht: (2024)
On the Importance of Multistability for Horizon Generalization in Reinforcement Learning
von: Bakija, Asad, et al.
Veröffentlicht: (2026)
von: Bakija, Asad, et al.
Veröffentlicht: (2026)
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks
von: Erba, Vittorio, et al.
Veröffentlicht: (2025)
von: Erba, Vittorio, et al.
Veröffentlicht: (2025)
Fundamental Limits of Matrix Sensing: Exact Asymptotics, Universality, and Applications
von: Xu, Yizhou, et al.
Veröffentlicht: (2025)
von: Xu, Yizhou, et al.
Veröffentlicht: (2025)
Overcoming the Challenges of Batch Normalization in Federated Learning
von: Guerraoui, Rachid, et al.
Veröffentlicht: (2024)
von: Guerraoui, Rachid, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions
von: Arnaboldi, Luca, et al.
Veröffentlicht: (2024) -
The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents
von: Dandi, Yatin, et al.
Veröffentlicht: (2024) -
How Two-Layer Neural Networks Learn, One (Giant) Step at a Time
von: Dandi, Yatin, et al.
Veröffentlicht: (2023) -
Deep Learning of Compositional Targets with Hierarchical Spectral Methods
von: Tabanelli, Hugo, et al.
Veröffentlicht: (2026) -
The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
von: Dandi, Yatin, et al.
Veröffentlicht: (2025)