Linear Recursive Feature Machines provably recover low-rank matrices
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Radhakrishnan, Adityanarayanan, Belkin, Mikhail, Drusvyatskiy, Dmitriy |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
par: Zhu, Libin, et autres
Publié: (2026)
par: Zhu, Libin, et autres
Publié: (2026)
Context-Scaling versus Task-Scaling in In-Context Learning
par: Abedsoltan, Amirhesam, et autres
Publié: (2024)
par: Abedsoltan, Amirhesam, et autres
Publié: (2024)
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
par: Beaglehole, Daniel, et autres
Publié: (2025)
par: Beaglehole, Daniel, et autres
Publié: (2025)
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
par: Zhu, Libin, et autres
Publié: (2023)
par: Zhu, Libin, et autres
Publié: (2023)
Quadratic models for understanding catapult dynamics of neural networks
par: Zhu, Libin, et autres
Publié: (2022)
par: Zhu, Libin, et autres
Publié: (2022)
Toward universal steering and monitoring of AI models
par: Beaglehole, Daniel, et autres
Publié: (2025)
par: Beaglehole, Daniel, et autres
Publié: (2025)
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
par: Mallinar, Neil, et autres
Publié: (2024)
par: Mallinar, Neil, et autres
Publié: (2024)
The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks
par: Cha, Taehun, et autres
Publié: (2026)
par: Cha, Taehun, et autres
Publié: (2026)
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
par: Davis, Damek, et autres
Publié: (2026)
par: Davis, Damek, et autres
Publié: (2026)
When do spectral gradient updates help in deep learning?
par: Davis, Damek, et autres
Publié: (2025)
par: Davis, Damek, et autres
Publié: (2025)
Contextual Linear Activation Steering of Language Models
par: Hsu, Brandon, et autres
Publié: (2026)
par: Hsu, Brandon, et autres
Publié: (2026)
Efficient and accurate steering of Large Language Models through attention-guided feature learning
par: Davarmanesh, Parmida, et autres
Publié: (2026)
par: Davarmanesh, Parmida, et autres
Publié: (2026)
High-dimensional Limit of SGD for Diagonal Linear Networks
par: Malaxechebarría, Begoña García, et autres
Publié: (2026)
par: Malaxechebarría, Begoña García, et autres
Publié: (2026)
Stochastic Approximation with Decision-Dependent Distributions: Asymptotic Normality and Optimality
par: Cutler, Joshua, et autres
Publié: (2022)
par: Cutler, Joshua, et autres
Publié: (2022)
Iteratively reweighted kernel machines efficiently learn sparse functions
par: Zhu, Libin, et autres
Publié: (2025)
par: Zhu, Libin, et autres
Publié: (2025)
Gradient descent with adaptive stepsize converges (nearly) linearly under fourth-order growth
par: Davis, Damek, et autres
Publié: (2024)
par: Davis, Damek, et autres
Publié: (2024)
Invariant Kernels: Rank Stabilization and Generalization Across Dimensions
par: Díaz, Mateo, et autres
Publié: (2025)
par: Díaz, Mateo, et autres
Publié: (2025)
The radius of statistical efficiency
par: Cutler, Joshua, et autres
Publié: (2024)
par: Cutler, Joshua, et autres
Publié: (2024)
On the Nystrom Approximation for Preconditioning in Kernel Machines
par: Abedsoltan, Amirhesam, et autres
Publié: (2023)
par: Abedsoltan, Amirhesam, et autres
Publié: (2023)
Breaking Data Symmetry is Needed For Generalization in Feature Learning Kernels
par: Bernal, Marcel Tomàs, et autres
Publié: (2026)
par: Bernal, Marcel Tomàs, et autres
Publié: (2026)
Shallow diffusion networks provably learn hidden low-dimensional structure
par: Boffi, Nicholas M., et autres
Publié: (2024)
par: Boffi, Nicholas M., et autres
Publié: (2024)
Online Covariance Estimation in Nonsmooth Stochastic Approximation
par: Jiang, Liwei, et autres
Publié: (2025)
par: Jiang, Liwei, et autres
Publié: (2025)
More is Better in Modern Machine Learning: when Infinite Overparameterization is Optimal and Overfitting is Obligatory
par: Simon, James B., et autres
Publié: (2023)
par: Simon, James B., et autres
Publié: (2023)
Computational and statistical lower bounds for low-rank estimation under general inhomogeneous noise
par: De, Debsurya, et autres
Publié: (2025)
par: De, Debsurya, et autres
Publié: (2025)
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
par: Boix-Adsera, Enric, et autres
Publié: (2025)
par: Boix-Adsera, Enric, et autres
Publié: (2025)
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
par: Mirtaheri, Parsa, et autres
Publié: (2026)
par: Mirtaheri, Parsa, et autres
Publié: (2026)
Entrywise error bounds for low-rank approximations of kernel matrices
par: Modell, Alexander
Publié: (2024)
par: Modell, Alexander
Publié: (2024)
On provable privacy vulnerabilities of graph representations
par: Wu, Ruofan, et autres
Publié: (2024)
par: Wu, Ruofan, et autres
Publié: (2024)
Mirror Descent on Reproducing Kernel Banach Spaces
par: Kumar, Akash, et autres
Publié: (2024)
par: Kumar, Akash, et autres
Publié: (2024)
General and Efficient Steering of Unconditional Diffusion
par: Wang, Qingsong, et autres
Publié: (2026)
par: Wang, Qingsong, et autres
Publié: (2026)
A Gap Between the Gaussian RKHS and Neural Networks: An Infinite-Center Asymptotic Analysis
par: Kumar, Akash, et autres
Publié: (2025)
par: Kumar, Akash, et autres
Publié: (2025)
When big data actually are low-rank, or entrywise approximation of certain function-generated matrices
par: Budzinskiy, Stanislav
Publié: (2024)
par: Budzinskiy, Stanislav
Publié: (2024)
Recursive State Inference for Linear PASFA
par: Rishi, Vishal
Publié: (2025)
par: Rishi, Vishal
Publié: (2025)
Fast training of large kernel models with delayed projections
par: Abedsoltan, Amirhesam, et autres
Publié: (2024)
par: Abedsoltan, Amirhesam, et autres
Publié: (2024)
Average gradient outer product as a mechanism for deep neural collapse
par: Beaglehole, Daniel, et autres
Publié: (2024)
par: Beaglehole, Daniel, et autres
Publié: (2024)
Localization from structured distance matrices via low-rank matrix recovery
par: Lichtenberg, Samuel, et autres
Publié: (2023)
par: Lichtenberg, Samuel, et autres
Publié: (2023)
Steering Autoregressive Music Generation with Recursive Feature Machines
par: Zhao, Daniel, et autres
Publié: (2025)
par: Zhao, Daniel, et autres
Publié: (2025)
Attention layers provably solve single-location regression
par: Marion, Pierre, et autres
Publié: (2024)
par: Marion, Pierre, et autres
Publié: (2024)
Learning efficient and provably convergent splitting methods
par: Kreusser, L. M., et autres
Publié: (2024)
par: Kreusser, L. M., et autres
Publié: (2024)
Amortised and provably-robust simulation-based inference
par: Bharti, Ayush, et autres
Publié: (2026)
par: Bharti, Ayush, et autres
Publié: (2026)
Documents similaires
-
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
par: Zhu, Libin, et autres
Publié: (2026) -
Context-Scaling versus Task-Scaling in In-Context Learning
par: Abedsoltan, Amirhesam, et autres
Publié: (2024) -
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
par: Beaglehole, Daniel, et autres
Publié: (2025) -
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
par: Zhu, Libin, et autres
Publié: (2023) -
Quadratic models for understanding catapult dynamics of neural networks
par: Zhu, Libin, et autres
Publié: (2022)