Fast training of large kernel models with delayed projections
Fuente:
arXiv
Saved in:
| Main Authors: | Abedsoltan, Amirhesam, Ma, Siyuan, Pandit, Parthe, Belkin, Mikhail |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Nystrom Approximation for Preconditioning in Kernel Machines
by: Abedsoltan, Amirhesam, et al.
Published: (2023)
by: Abedsoltan, Amirhesam, et al.
Published: (2023)
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
by: Mallinar, Neil, et al.
Published: (2022)
by: Mallinar, Neil, et al.
Published: (2022)
Context-Scaling versus Task-Scaling in In-Context Learning
by: Abedsoltan, Amirhesam, et al.
Published: (2024)
by: Abedsoltan, Amirhesam, et al.
Published: (2024)
Mirror Descent on Reproducing Kernel Banach Spaces
by: Kumar, Akash, et al.
Published: (2024)
by: Kumar, Akash, et al.
Published: (2024)
Feature maps for the Laplacian kernel and its generalizations
by: Ahir, Sudhendu, et al.
Published: (2025)
by: Ahir, Sudhendu, et al.
Published: (2025)
Task Generalization With AutoRegressive Compositional Structure: Can Learning From $D$ Tasks Generalize to $D^{T}$ Tasks?
by: Abedsoltan, Amirhesam, et al.
Published: (2025)
by: Abedsoltan, Amirhesam, et al.
Published: (2025)
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
by: Mallinar, Neil, et al.
Published: (2024)
by: Mallinar, Neil, et al.
Published: (2024)
Asymptotic convexity of wide and shallow neural networks
by: Borkar, Vivek, et al.
Published: (2025)
by: Borkar, Vivek, et al.
Published: (2025)
Universality of Kernel Random Matrices and Kernel Regression in the Quadratic Regime
by: Pandit, Parthe, et al.
Published: (2024)
by: Pandit, Parthe, et al.
Published: (2024)
Eigenvectors of the De Bruijn Graph Laplacian: A Natural Basis for the Cut and Cycle Space
by: Philippakis, Anthony, et al.
Published: (2024)
by: Philippakis, Anthony, et al.
Published: (2024)
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
by: Zhu, Libin, et al.
Published: (2023)
by: Zhu, Libin, et al.
Published: (2023)
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
by: Beaglehole, Daniel, et al.
Published: (2025)
by: Beaglehole, Daniel, et al.
Published: (2025)
Linear Recursive Feature Machines provably recover low-rank matrices
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
by: Mirtaheri, Parsa, et al.
Published: (2026)
by: Mirtaheri, Parsa, et al.
Published: (2026)
Quadratic models for understanding catapult dynamics of neural networks
by: Zhu, Libin, et al.
Published: (2022)
by: Zhu, Libin, et al.
Published: (2022)
On-Policy Distillation of Language Models for Autonomous Vehicle Motion Planning
by: Afsharrad, Amirhossein, et al.
Published: (2026)
by: Afsharrad, Amirhossein, et al.
Published: (2026)
General and Efficient Steering of Unconditional Diffusion
by: Wang, Qingsong, et al.
Published: (2026)
by: Wang, Qingsong, et al.
Published: (2026)
A Gap Between the Gaussian RKHS and Neural Networks: An Infinite-Center Asymptotic Analysis
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
Average gradient outer product as a mechanism for deep neural collapse
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Breaking Data Symmetry is Needed For Generalization in Feature Learning Kernels
by: Bernal, Marcel Tomàs, et al.
Published: (2026)
by: Bernal, Marcel Tomàs, et al.
Published: (2026)
Toward universal steering and monitoring of AI models
by: Beaglehole, Daniel, et al.
Published: (2025)
by: Beaglehole, Daniel, et al.
Published: (2025)
More is Better in Modern Machine Learning: when Infinite Overparameterization is Optimal and Overfitting is Obligatory
by: Simon, James B., et al.
Published: (2023)
by: Simon, James B., et al.
Published: (2023)
Fast kernel methods: Sobolev, physics-informed, and additive models
by: Doumèche, Nathan, et al.
Published: (2025)
by: Doumèche, Nathan, et al.
Published: (2025)
Optimal differentially private kernel learning with random projection
by: Lee, Bonwoo, et al.
Published: (2025)
by: Lee, Bonwoo, et al.
Published: (2025)
The phase diagram of kernel interpolation in large dimensions
by: Zhang, Haobo, et al.
Published: (2024)
by: Zhang, Haobo, et al.
Published: (2024)
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
by: Boix-Adsera, Enric, et al.
Published: (2025)
by: Boix-Adsera, Enric, et al.
Published: (2025)
A discrete physics-informed training for projection-based reduced order models with neural networks
by: Sibuet, N., et al.
Published: (2025)
by: Sibuet, N., et al.
Published: (2025)
Benchmarking quantum machine learning kernel training for classification tasks
by: Alvarez-Estevez, Diego
Published: (2024)
by: Alvarez-Estevez, Diego
Published: (2024)
Snacks: a fast large-scale kernel SVM solver
by: Tanji, Sofiane, et al.
Published: (2023)
by: Tanji, Sofiane, et al.
Published: (2023)
Fast Spectrum Estimation of Some Kernel Matrices
by: Lepilov, Mikhail
Published: (2024)
by: Lepilov, Mikhail
Published: (2024)
Fast constrained sampling in pre-trained diffusion models
by: Graikos, Alexandros, et al.
Published: (2024)
by: Graikos, Alexandros, et al.
Published: (2024)
A fast and effective kernel two-sample test for large-scale data
by: Song, Hoseung, et al.
Published: (2021)
by: Song, Hoseung, et al.
Published: (2021)
ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
by: Ahrens, Lara, et al.
Published: (2025)
by: Ahrens, Lara, et al.
Published: (2025)
Post-training makes large language models less human-like
by: Binz, Marcel, et al.
Published: (2026)
by: Binz, Marcel, et al.
Published: (2026)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
by: Fu, Deqing, et al.
Published: (2026)
by: Fu, Deqing, et al.
Published: (2026)
Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers
by: Chertkov, Andrei, et al.
Published: (2025)
by: Chertkov, Andrei, et al.
Published: (2025)
Generalized vec trick for fast learning of pairwise kernel models
by: Viljanen, Markus, et al.
Published: (2020)
by: Viljanen, Markus, et al.
Published: (2020)
Finetuning greedy kernel models by exchange algorithms
by: Wenzel, Tizian, et al.
Published: (2024)
by: Wenzel, Tizian, et al.
Published: (2024)
Hybrid model of the kernel method for quantum computers
by: de Borba, Jhordan Silveira, et al.
Published: (2024)
by: de Borba, Jhordan Silveira, et al.
Published: (2024)
Enforcing governing equation constraints in neural PDE solvers via training-free projections
by: Rochman, Omer, et al.
Published: (2025)
by: Rochman, Omer, et al.
Published: (2025)
Similar Items
-
On the Nystrom Approximation for Preconditioning in Kernel Machines
by: Abedsoltan, Amirhesam, et al.
Published: (2023) -
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
by: Mallinar, Neil, et al.
Published: (2022) -
Context-Scaling versus Task-Scaling in In-Context Learning
by: Abedsoltan, Amirhesam, et al.
Published: (2024) -
Mirror Descent on Reproducing Kernel Banach Spaces
by: Kumar, Akash, et al.
Published: (2024) -
Feature maps for the Laplacian kernel and its generalizations
by: Ahir, Sudhendu, et al.
Published: (2025)