Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
Fuente:
arXiv
Saved in:
| Main Authors: | Njaradi, Valentina, Dominé, Clémentine, Swanson, Rachel, Mondelli, Marco, Saxe, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
by: Anguita, Nicolas, et al.
Published: (2026)
by: Anguita, Nicolas, et al.
Published: (2026)
Optimal Learning Rate Schedule for Balancing Effort and Performance
by: Njaradi, Valentina, et al.
Published: (2026)
by: Njaradi, Valentina, et al.
Published: (2026)
High-Dimensional Private Linear Regression with Optimal Rates
by: Bombari, Simone, et al.
Published: (2025)
by: Bombari, Simone, et al.
Published: (2025)
Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-Parameterization
by: Bombari, Simone, et al.
Published: (2025)
by: Bombari, Simone, et al.
Published: (2025)
A Theory of Initialisation's Impact on Specialisation
by: Jarvis, Devon, et al.
Published: (2025)
by: Jarvis, Devon, et al.
Published: (2025)
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
by: Dominé, Clémentine C. J., et al.
Published: (2024)
by: Dominé, Clémentine C. J., et al.
Published: (2024)
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
by: Kunin, Daniel, et al.
Published: (2024)
by: Kunin, Daniel, et al.
Published: (2024)
Optimal Regularization for Performative Learning
by: Cyffers, Edwige, et al.
Published: (2025)
by: Cyffers, Edwige, et al.
Published: (2025)
How Spurious Features Are Memorized: Precise Analysis for Random and NTK Features
by: Bombari, Simone, et al.
Published: (2023)
by: Bombari, Simone, et al.
Published: (2023)
High-dimensional Analysis of Synthetic Data Selection
by: Rezaei, Parham, et al.
Published: (2025)
by: Rezaei, Parham, et al.
Published: (2025)
Matrix Denoising with Doubly Heteroscedastic Noise: Fundamental Limits and Optimal Spectral Methods
by: Zhang, Yihan, et al.
Published: (2024)
by: Zhang, Yihan, et al.
Published: (2024)
When Representations Align: Universality in Representation Learning Dynamics
by: van Rossem, Loek, et al.
Published: (2024)
by: van Rossem, Loek, et al.
Published: (2024)
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
by: Súkeník, Peter, et al.
Published: (2024)
by: Súkeník, Peter, et al.
Published: (2024)
Optimal Estimation in Orthogonally Invariant Generalized Linear Models: Spectral Initialization and Approximate Message Passing
by: Zhang, Yihan, et al.
Published: (2026)
by: Zhang, Yihan, et al.
Published: (2026)
Precise Asymptotics for Spectral Methods in Mixed Generalized Linear Models
by: Zhang, Yihan, et al.
Published: (2022)
by: Zhang, Yihan, et al.
Published: (2022)
Understanding Unimodal Bias in Multimodal Deep Linear Networks
by: Zhang, Yedi, et al.
Published: (2023)
by: Zhang, Yedi, et al.
Published: (2023)
Neural Collapse Beyond the Unconstrained Features Model: Landscape, Dynamics, and Generalization in the Mean-Field Regime
by: Wu, Diyuan, et al.
Published: (2025)
by: Wu, Diyuan, et al.
Published: (2025)
Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers
by: Súkeník, Peter, et al.
Published: (2025)
by: Súkeník, Peter, et al.
Published: (2025)
When Are Bias-Free ReLU Networks Effectively Linear Networks?
by: Zhang, Yedi, et al.
Published: (2024)
by: Zhang, Yedi, et al.
Published: (2024)
Spectral Estimators for Multi-Index Models: Precise Asymptotics and Optimal Weak Recovery
by: Kovačević, Filip, et al.
Published: (2025)
by: Kovačević, Filip, et al.
Published: (2025)
Privacy for Free in the Overparameterized Regime
by: Bombari, Simone, et al.
Published: (2024)
by: Bombari, Simone, et al.
Published: (2024)
Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features
by: Bombari, Simone, et al.
Published: (2024)
by: Bombari, Simone, et al.
Published: (2024)
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
by: Dragutinović, Sara, et al.
Published: (2025)
by: Dragutinović, Sara, et al.
Published: (2025)
Training Dynamics of In-Context Learning in Linear Attention
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
High-dimensional Analysis of Knowledge Distillation: Weak-to-Strong Generalization and Scaling Laws
by: Ildiz, M. Emrullah, et al.
Published: (2024)
by: Ildiz, M. Emrullah, et al.
Published: (2024)
Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions
by: Lee, Andrew, et al.
Published: (2026)
by: Lee, Andrew, et al.
Published: (2026)
Spectral Estimators for Structured Generalized Linear Models via Approximate Message Passing
by: Zhang, Yihan, et al.
Published: (2023)
by: Zhang, Yihan, et al.
Published: (2023)
Improved Convergence of Score-Based Diffusion Models via Prediction-Correction
by: Pedrotti, Francesco, et al.
Published: (2023)
by: Pedrotti, Francesco, et al.
Published: (2023)
Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression
by: Wu, Diyuan, et al.
Published: (2026)
by: Wu, Diyuan, et al.
Published: (2026)
A Law of Data Reconstruction for Random Features (and Beyond)
by: Iurada, Leonardo, et al.
Published: (2025)
by: Iurada, Leonardo, et al.
Published: (2025)
Average gradient outer product as a mechanism for deep neural collapse
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Linearized Optimal Transport for Analysis of High-Dimensional Point-Cloud and Single-Cell Data
by: Wang, Tianxiang, et al.
Published: (2025)
by: Wang, Tianxiang, et al.
Published: (2025)
Feature Clock: High-Dimensional Effects in Two-Dimensional Plots
by: Ovcharenko, Olga, et al.
Published: (2024)
by: Ovcharenko, Olga, et al.
Published: (2024)
Minimax Rate-Optimal Algorithms for High-Dimensional Stochastic Linear Bandits
by: Liu, Jingyu, et al.
Published: (2025)
by: Liu, Jingyu, et al.
Published: (2025)
Nonlinear dynamics of localization in neural receptive fields
by: Lufkin, Leon, et al.
Published: (2025)
by: Lufkin, Leon, et al.
Published: (2025)
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Algorithm Development in Neural Networks: Insights from the Streaming Parity Task
by: van Rossem, Loek, et al.
Published: (2025)
by: van Rossem, Loek, et al.
Published: (2025)
Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
by: Nam, Yoonsoo, et al.
Published: (2025)
by: Nam, Yoonsoo, et al.
Published: (2025)
Attention with Trained Embeddings Provably Selects Important Tokens
by: Wu, Diyuan, et al.
Published: (2025)
by: Wu, Diyuan, et al.
Published: (2025)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
by: Jacot, Arthur, et al.
Published: (2024)
by: Jacot, Arthur, et al.
Published: (2024)
Similar Items
-
A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
by: Anguita, Nicolas, et al.
Published: (2026) -
Optimal Learning Rate Schedule for Balancing Effort and Performance
by: Njaradi, Valentina, et al.
Published: (2026) -
High-Dimensional Private Linear Regression with Optimal Rates
by: Bombari, Simone, et al.
Published: (2025) -
Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-Parameterization
by: Bombari, Simone, et al.
Published: (2025) -
A Theory of Initialisation's Impact on Specialisation
by: Jarvis, Devon, et al.
Published: (2025)