The Power of Random Features and the Limits of Distribution-Free Gradient Descent
Fuente:
arXiv
Saved in:
| Main Authors: | Karchmer, Ari, Malach, Eran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Stronger Computational Separations Between Multimodal and Unimodal Machine Learning
by: Karchmer, Ari
Published: (2024)
by: Karchmer, Ari
Published: (2024)
Auto-Regressive Next-Token Predictors are Universal Learners
by: Malach, Eran
Published: (2023)
by: Malach, Eran
Published: (2023)
Efficiently Verifiable Proofs of Data Attribution
by: Karchmer, Ari, et al.
Published: (2025)
by: Karchmer, Ari, et al.
Published: (2025)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
LLM Priors for ERM over Programs
by: Singhal, Shivam, et al.
Published: (2025)
by: Singhal, Shivam, et al.
Published: (2025)
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
by: Tsilivis, Nikolaos, et al.
Published: (2025)
by: Tsilivis, Nikolaos, et al.
Published: (2025)
Covert Quantum Learning: Privately and Verifiably Learning from Quantum Data
by: Anand, Abhishek, et al.
Published: (2025)
by: Anand, Abhishek, et al.
Published: (2025)
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024)
by: Hou, Kaiying, et al.
Published: (2024)
Distributed Gradient Descent for Functional Learning
by: Yu, Zhan, et al.
Published: (2023)
by: Yu, Zhan, et al.
Published: (2023)
Randomness and Interpolation Improve Gradient Descent
by: Li, Jiawen, et al.
Published: (2025)
by: Li, Jiawen, et al.
Published: (2025)
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
by: Edelman, Benjamin L., et al.
Published: (2024)
by: Edelman, Benjamin L., et al.
Published: (2024)
Automated Feature Labeling with Token-Space Gradient Descent
by: Schulz, Julian, et al.
Published: (2025)
by: Schulz, Julian, et al.
Published: (2025)
Limit Theorems for Stochastic Gradient Descent with Infinite Variance
by: Blanchet, Jose, et al.
Published: (2024)
by: Blanchet, Jose, et al.
Published: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Central Limit Theorems for Stochastic Gradient Descent Quantile Estimators
by: Wei, Ziyang, et al.
Published: (2025)
by: Wei, Ziyang, et al.
Published: (2025)
Functional Central Limit Theorem for Stochastic Gradient Descent
by: Flamand, Kessang, et al.
Published: (2026)
by: Flamand, Kessang, et al.
Published: (2026)
Don't Stop Me Now: Embedding Based Scheduling for LLMs
by: Shahout, Rana, et al.
Published: (2024)
by: Shahout, Rana, et al.
Published: (2024)
Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
Inversion-Free Natural Gradient Descent on Riemannian Manifolds
by: Draca, Dario, et al.
Published: (2026)
by: Draca, Dario, et al.
Published: (2026)
The Limit Points of (Optimistic) Gradient Descent in Min-Max Optimization
by: Daskalakis, Constantinos, et al.
Published: (2018)
by: Daskalakis, Constantinos, et al.
Published: (2018)
Limited Memory Online Gradient Descent for Kernelized Pairwise Learning with Dynamic Averaging
by: AlQuabeh, Hilal, et al.
Published: (2024)
by: AlQuabeh, Hilal, et al.
Published: (2024)
Occam Gradient Descent
by: Kausik, B. N.
Published: (2024)
by: Kausik, B. N.
Published: (2024)
On Double Descent in Reinforcement Learning with LSTD and Random Features
by: Brellmann, David, et al.
Published: (2023)
by: Brellmann, David, et al.
Published: (2023)
Noise Balance and Stationary Distribution of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2023)
by: Ziyin, Liu, et al.
Published: (2023)
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024)
by: Morwani, Depen, et al.
Published: (2024)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024)
by: Prabhakar, Akshara, et al.
Published: (2024)
Multiple Wasserstein Gradient Descent Algorithm for Multi-Objective Distributional Optimization
by: Nguyen, Dai Hai, et al.
Published: (2025)
by: Nguyen, Dai Hai, et al.
Published: (2025)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
A Mirror Descent-Based Algorithm for Corruption-Tolerant Distributed Gradient Descent
by: Wang, Shuche, et al.
Published: (2024)
by: Wang, Shuche, et al.
Published: (2024)
Accelerating Natural Gradient Descent for PINNs with Randomized Numerical Linear Algebra
by: Bioli, Ivan, et al.
Published: (2025)
by: Bioli, Ivan, et al.
Published: (2025)
Rank-1 Matrix Completion with Gradient Descent and Small Random Initialization
by: Kim, Daesung, et al.
Published: (2022)
by: Kim, Daesung, et al.
Published: (2022)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
A Stein Gradient Descent Approach for Doubly Intractable Distributions
by: Lee, Heesang, et al.
Published: (2024)
by: Lee, Heesang, et al.
Published: (2024)
Stacking as Accelerated Gradient Descent
by: Agarwal, Naman, et al.
Published: (2024)
by: Agarwal, Naman, et al.
Published: (2024)
Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks
by: Li, Binghui, et al.
Published: (2024)
by: Li, Binghui, et al.
Published: (2024)
Stochastic Adaptive Gradient Descent Without Descent
by: Aujol, Jean-François, et al.
Published: (2025)
by: Aujol, Jean-François, et al.
Published: (2025)
Random Matrix Theory for Stochastic Gradient Descent
by: Park, Chanju, et al.
Published: (2024)
by: Park, Chanju, et al.
Published: (2024)
Corner Gradient Descent
by: Yarotsky, Dmitry
Published: (2025)
by: Yarotsky, Dmitry
Published: (2025)
Similar Items
-
On Stronger Computational Separations Between Multimodal and Unimodal Machine Learning
by: Karchmer, Ari
Published: (2024) -
Auto-Regressive Next-Token Predictors are Universal Learners
by: Malach, Eran
Published: (2023) -
Efficiently Verifiable Proofs of Data Attribution
by: Karchmer, Ari, et al.
Published: (2025) -
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
by: Qin, Tian, et al.
Published: (2025) -
LLM Priors for ERM over Programs
by: Singhal, Shivam, et al.
Published: (2025)