A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Anguita, Nicolas, Locatello, Francesco, Saxe, Andrew M., Mondelli, Marco, Mancini, Flavia, Lippl, Samuel, Domine, Clementine |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
by: Njaradi, Valentina, et al.
Published: (2026)
by: Njaradi, Valentina, et al.
Published: (2026)
A Theory of Initialisation's Impact on Specialisation
by: Jarvis, Devon, et al.
Published: (2025)
by: Jarvis, Devon, et al.
Published: (2025)
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
by: Dominé, Clémentine C. J., et al.
Published: (2024)
by: Dominé, Clémentine C. J., et al.
Published: (2024)
High-dimensional Analysis of Synthetic Data Selection
by: Rezaei, Parham, et al.
Published: (2025)
by: Rezaei, Parham, et al.
Published: (2025)
Inductive biases of multi-task learning and finetuning: multiple regimes of feature reuse
by: Lippl, Samuel, et al.
Published: (2023)
by: Lippl, Samuel, et al.
Published: (2023)
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
by: Kunin, Daniel, et al.
Published: (2024)
by: Kunin, Daniel, et al.
Published: (2024)
Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-Parameterization
by: Bombari, Simone, et al.
Published: (2025)
by: Bombari, Simone, et al.
Published: (2025)
How Spurious Features Are Memorized: Precise Analysis for Random and NTK Features
by: Bombari, Simone, et al.
Published: (2023)
by: Bombari, Simone, et al.
Published: (2023)
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
by: Súkeník, Peter, et al.
Published: (2024)
by: Súkeník, Peter, et al.
Published: (2024)
Causal Learning with the Invariance Principle
by: Montagna, Francesco, et al.
Published: (2026)
by: Montagna, Francesco, et al.
Published: (2026)
Understanding Unimodal Bias in Multimodal Deep Linear Networks
by: Zhang, Yedi, et al.
Published: (2023)
by: Zhang, Yedi, et al.
Published: (2023)
A mathematical theory of balancing relational generalization and memorization
by: Cheng, Luke, et al.
Published: (2026)
by: Cheng, Luke, et al.
Published: (2026)
When does compositional structure yield compositional generalization? A kernel theory
by: Lippl, Samuel, et al.
Published: (2024)
by: Lippl, Samuel, et al.
Published: (2024)
When Are Bias-Free ReLU Networks Effectively Linear Networks?
by: Zhang, Yedi, et al.
Published: (2024)
by: Zhang, Yedi, et al.
Published: (2024)
Improved Convergence of Score-Based Diffusion Models via Prediction-Correction
by: Pedrotti, Francesco, et al.
Published: (2023)
by: Pedrotti, Francesco, et al.
Published: (2023)
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Controlling Transient Amplification Improves Long-horizon Rollouts
by: Pervez, Adeel, et al.
Published: (2026)
by: Pervez, Adeel, et al.
Published: (2026)
Adjusting Pretrained Backbones for Performativity
by: Demirel, Berker, et al.
Published: (2024)
by: Demirel, Berker, et al.
Published: (2024)
Neural Collapse Beyond the Unconstrained Features Model: Landscape, Dynamics, and Generalization in the Mean-Field Regime
by: Wu, Diyuan, et al.
Published: (2025)
by: Wu, Diyuan, et al.
Published: (2025)
The Rate-Distortion-Polysemanticity Tradeoff in SAEs
by: Mencattini, Tommaso, et al.
Published: (2026)
by: Mencattini, Tommaso, et al.
Published: (2026)
Out-of-Distribution Detection with Relative Angles
by: Demirel, Berker, et al.
Published: (2024)
by: Demirel, Berker, et al.
Published: (2024)
Statistical and structural identifiability in representation learning
by: Nelson, Walter, et al.
Published: (2026)
by: Nelson, Walter, et al.
Published: (2026)
Navigating the Latent Space Dynamics of Neural Models
by: Fumero, Marco, et al.
Published: (2025)
by: Fumero, Marco, et al.
Published: (2025)
Privacy for Free in the Overparameterized Regime
by: Bombari, Simone, et al.
Published: (2024)
by: Bombari, Simone, et al.
Published: (2024)
Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features
by: Bombari, Simone, et al.
Published: (2024)
by: Bombari, Simone, et al.
Published: (2024)
Towards a holistic understanding of Selection Bias for Causal Effect Identification
by: Qiu, Yiwen, et al.
Published: (2026)
by: Qiu, Yiwen, et al.
Published: (2026)
Latent Functional Maps: a spectral framework for representation alignment
by: Fumero, Marco, et al.
Published: (2024)
by: Fumero, Marco, et al.
Published: (2024)
Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks
by: Jarvis, Devon, et al.
Published: (2025)
by: Jarvis, Devon, et al.
Published: (2025)
Shaping Inductive Bias in Diffusion Models through Frequency-Based Noise Control
by: Jiralerspong, Thomas, et al.
Published: (2025)
by: Jiralerspong, Thomas, et al.
Published: (2025)
Addressing Instrument-Outcome Confounding in Mendelian Randomization through Representation Learning
by: Huang, Shimeng, et al.
Published: (2026)
by: Huang, Shimeng, et al.
Published: (2026)
Toward Identifiable Sparse Autoencoders
by: Nelson, Walter, et al.
Published: (2026)
by: Nelson, Walter, et al.
Published: (2026)
Marrying Causal Representation Learning with Dynamical Systems for Science
by: Yao, Dingling, et al.
Published: (2024)
by: Yao, Dingling, et al.
Published: (2024)
Mechanistic PDE Networks for Discovery of Governing Equations
by: Pervez, Adeel, et al.
Published: (2025)
by: Pervez, Adeel, et al.
Published: (2025)
Learning Discrete Diffusion of Graphs via Free-Energy Gradient Flows
by: Rancati, Dario, et al.
Published: (2026)
by: Rancati, Dario, et al.
Published: (2026)
Optimal Regularization for Performative Learning
by: Cyffers, Edwige, et al.
Published: (2025)
by: Cyffers, Edwige, et al.
Published: (2025)
On the Role of Inductive Bias in Time-Series Pretraining: A Case Study in Learning Generalizable Representations for Clinical Time Series
by: Dey, Sharmita, et al.
Published: (2026)
by: Dey, Sharmita, et al.
Published: (2026)
Soft Geometric Inductive Bias for Object Centric Dynamics
by: Linander, Hampus, et al.
Published: (2025)
by: Linander, Hampus, et al.
Published: (2025)
Dataset Difficulty and the Role of Inductive Bias
by: Kwok, Devin, et al.
Published: (2024)
by: Kwok, Devin, et al.
Published: (2024)
Interpolated-MLPs: Controllable Inductive Bias
by: Wu, Sean, et al.
Published: (2024)
by: Wu, Sean, et al.
Published: (2024)
Towards Exact Computation of Inductive Bias
by: Boopathy, Akhilan, et al.
Published: (2024)
by: Boopathy, Akhilan, et al.
Published: (2024)
Similar Items
-
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
by: Njaradi, Valentina, et al.
Published: (2026) -
A Theory of Initialisation's Impact on Specialisation
by: Jarvis, Devon, et al.
Published: (2025) -
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
by: Dominé, Clémentine C. J., et al.
Published: (2024) -
High-dimensional Analysis of Synthetic Data Selection
by: Rezaei, Parham, et al.
Published: (2025) -
Inductive biases of multi-task learning and finetuning: multiple regimes of feature reuse
by: Lippl, Samuel, et al.
Published: (2023)