The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
Fuente:
arXiv
Saved in:
| Main Authors: | Boix-Adsera, Enric, Mallinar, Neil, Simon, James B., Belkin, Mikhail |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Secret mixtures of experts inside your LLM
by: Boix-Adsera, Enric
Published: (2025)
by: Boix-Adsera, Enric
Published: (2025)
On the inductive bias of infinite-depth ResNets and the bottleneck rank
by: Boix-Adsera, Enric
Published: (2025)
by: Boix-Adsera, Enric
Published: (2025)
Toward universal steering and monitoring of AI models
by: Beaglehole, Daniel, et al.
Published: (2025)
by: Beaglehole, Daniel, et al.
Published: (2025)
Towards a theory of model distillation
by: Boix-Adsera, Enric
Published: (2024)
by: Boix-Adsera, Enric
Published: (2024)
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
by: Boix-Adsera, Enric, et al.
Published: (2025)
by: Boix-Adsera, Enric, et al.
Published: (2025)
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
by: Mallinar, Neil, et al.
Published: (2022)
by: Mallinar, Neil, et al.
Published: (2022)
Breaking Data Symmetry is Needed For Generalization in Feature Learning Kernels
by: Bernal, Marcel Tomàs, et al.
Published: (2026)
by: Bernal, Marcel Tomàs, et al.
Published: (2026)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
When can transformers reason with abstract symbols?
by: Boix-Adsera, Enric, et al.
Published: (2023)
by: Boix-Adsera, Enric, et al.
Published: (2023)
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
by: Abbe, Emmanuel, et al.
Published: (2022)
by: Abbe, Emmanuel, et al.
Published: (2022)
Feature contamination: Neural networks learn uncorrelated features and fail to generalize
by: Zhang, Tianren, et al.
Published: (2024)
by: Zhang, Tianren, et al.
Published: (2024)
A Gap Between the Gaussian RKHS and Neural Networks: An Infinite-Center Asymptotic Analysis
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
Impossibility Theorems for Feature Attribution
by: Bilodeau, Blair, et al.
Published: (2022)
by: Bilodeau, Blair, et al.
Published: (2022)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
by: Fu, Deqing, et al.
Published: (2026)
by: Fu, Deqing, et al.
Published: (2026)
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
by: Mirtaheri, Parsa, et al.
Published: (2026)
by: Mirtaheri, Parsa, et al.
Published: (2026)
General and Efficient Steering of Unconditional Diffusion
by: Wang, Qingsong, et al.
Published: (2026)
by: Wang, Qingsong, et al.
Published: (2026)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Comparative Analysis of QNN Architectures for Wind Power Prediction: Feature Maps and Ansatz Configurations
by: Hangun, Batuhan, et al.
Published: (2025)
by: Hangun, Batuhan, et al.
Published: (2025)
Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features
by: Wang, Jiuqi, et al.
Published: (2024)
by: Wang, Jiuqi, et al.
Published: (2024)
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
by: Mallinar, Neil, et al.
Published: (2024)
by: Mallinar, Neil, et al.
Published: (2024)
Operator Feature Neural Network for Symbolic Regression
by: Deng, Yusong, et al.
Published: (2024)
by: Deng, Yusong, et al.
Published: (2024)
Addressing divergent representations from causal interventions on neural networks
by: Grant, Satchel, et al.
Published: (2025)
by: Grant, Satchel, et al.
Published: (2025)
Unveiling the Power of Sparse Neural Networks for Feature Selection
by: Atashgahi, Zahra, et al.
Published: (2024)
by: Atashgahi, Zahra, et al.
Published: (2024)
Explicit Feature Interaction-aware Graph Neural Networks
by: Kim, Minkyu, et al.
Published: (2022)
by: Kim, Minkyu, et al.
Published: (2022)
InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders
by: Simon, Elana, et al.
Published: (2024)
by: Simon, Elana, et al.
Published: (2024)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
by: Chen, Zixiang, et al.
Published: (2025)
by: Chen, Zixiang, et al.
Published: (2025)
Training-free Graph Neural Networks and the Power of Labels as Features
by: Sato, Ryoma
Published: (2024)
by: Sato, Ryoma
Published: (2024)
Graph Neural Networks with Feature and Structure Aware Random Walk
by: Zhuo, Wei, et al.
Published: (2021)
by: Zhuo, Wei, et al.
Published: (2021)
Feature Learning Dynamics in Infinite-Depth Neural Networks
by: Yao, Zihan, et al.
Published: (2025)
by: Yao, Zihan, et al.
Published: (2025)
Mitigating Overfitting in Graph Neural Networks via Feature and Hyperplane Perturbation
by: Choi, Yoonhyuk, et al.
Published: (2022)
by: Choi, Yoonhyuk, et al.
Published: (2022)
TinyGraph: Joint Feature and Node Condensation for Graph Neural Networks
by: Liu, Yezi, et al.
Published: (2024)
by: Liu, Yezi, et al.
Published: (2024)
A Stable Neural Statistical Dependence Estimator for Autoencoder Feature Analysis
by: Hu, Bo, et al.
Published: (2026)
by: Hu, Bo, et al.
Published: (2026)
TangledFeatures: Robust Feature Selection in Highly Correlated Spaces
by: Sunny, Allen Daniel
Published: (2025)
by: Sunny, Allen Daniel
Published: (2025)
Augmenting learning in neuro-embodied systems through neurobiological first principles
by: Rodriguez-Garcia, Alejandro, et al.
Published: (2024)
by: Rodriguez-Garcia, Alejandro, et al.
Published: (2024)
Eigenvectors of the De Bruijn Graph Laplacian: A Natural Basis for the Cut and Cycle Space
by: Philippakis, Anthony, et al.
Published: (2024)
by: Philippakis, Anthony, et al.
Published: (2024)
Curriculum reinforcement learning with measurable task representation learning
by: Wen, Yongyan, et al.
Published: (2026)
by: Wen, Yongyan, et al.
Published: (2026)
Multi-View Graph Feature Propagation for Privacy Preservation and Feature Sparsity
by: Harari, Etzion, et al.
Published: (2025)
by: Harari, Etzion, et al.
Published: (2025)
Meta-learning how to Share Credit among Macro-Actions
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
Decoding complexity: how machine learning is redefining scientific discovery
by: Vinuesa, Ricardo, et al.
Published: (2024)
by: Vinuesa, Ricardo, et al.
Published: (2024)
The Impact of Feature Embedding Placement in the Ansatz of a Quantum Kernel in QSVMs
by: Salmenperä, Ilmo, et al.
Published: (2024)
by: Salmenperä, Ilmo, et al.
Published: (2024)
Similar Items
-
Secret mixtures of experts inside your LLM
by: Boix-Adsera, Enric
Published: (2025) -
On the inductive bias of infinite-depth ResNets and the bottleneck rank
by: Boix-Adsera, Enric
Published: (2025) -
Toward universal steering and monitoring of AI models
by: Beaglehole, Daniel, et al.
Published: (2025) -
Towards a theory of model distillation
by: Boix-Adsera, Enric
Published: (2024) -
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
by: Boix-Adsera, Enric, et al.
Published: (2025)