Gradient Descent Fails to Learn High-frequency Functions and Modular Arithmetic
Fuente:
arXiv
Saved in:
| Main Authors: | Takhanov, Rustem, Tezekbayev, Maxat, Pak, Artur, Bolatov, Arman, Assylbekov, Zhenisbek |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Linear Discriminant Analysis Revisited
by: Tezekbayev, Maxat, et al.
Published: (2026)
by: Tezekbayev, Maxat, et al.
Published: (2026)
Simplex Deep Linear Discriminant Analysis
by: Tezekbayev, Maxat, et al.
Published: (2026)
by: Tezekbayev, Maxat, et al.
Published: (2026)
Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel Thresholding
by: Takhanov, Rustem, et al.
Published: (2026)
by: Takhanov, Rustem, et al.
Published: (2026)
Overspecified Mixture Discriminant Analysis: Exponential Convergence, Statistical Guarantees, and Remote Sensing Applications
by: Bolatov, Arman, et al.
Published: (2025)
by: Bolatov, Arman, et al.
Published: (2025)
Classifier Performance on Long‐Tail Distributions
by: Artur Pak, et al.
Published: (2025)
by: Artur Pak, et al.
Published: (2025)
Learning Overspecified Gaussian Mixtures Exponentially Fast with the EM Algorithm
by: Assylbekov, Zhenisbek, et al.
Published: (2025)
by: Assylbekov, Zhenisbek, et al.
Published: (2025)
On the Intrinsic Dimensions of Data in Kernel Learning
by: Takhanov, Rustem
Published: (2026)
by: Takhanov, Rustem
Published: (2026)
The informativeness of the gradient revisited
by: Takhanov, Rustem
Published: (2025)
by: Takhanov, Rustem
Published: (2025)
Multi-layer random features and the approximation power of neural networks
by: Takhanov, Rustem
Published: (2024)
by: Takhanov, Rustem
Published: (2024)
Non-asymptotic spectral bounds on the $\varepsilon$-entropy of kernel classes
by: Takhanov, Rustem
Published: (2022)
by: Takhanov, Rustem
Published: (2022)
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
by: Bolatov, Arman, et al.
Published: (2026)
by: Bolatov, Arman, et al.
Published: (2026)
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
Byzantine-Robust Optimization under $(L_0, L_1)$-Smoothness
by: Bolatov, Arman, et al.
Published: (2026)
by: Bolatov, Arman, et al.
Published: (2026)
Distributed Gradient Descent for Functional Learning
by: Yu, Zhan, et al.
Published: (2023)
by: Yu, Zhan, et al.
Published: (2023)
Learning High-Dimensional Parity Functions with Product Networks using Gradient Descent
by: Larue, Guillaume, et al.
Published: (2026)
by: Larue, Guillaume, et al.
Published: (2026)
The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
by: Dandi, Yatin, et al.
Published: (2025)
by: Dandi, Yatin, et al.
Published: (2025)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026)
by: Riabinin, Artem, et al.
Published: (2026)
Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
by: Cheng, Xiang, et al.
Published: (2023)
by: Cheng, Xiang, et al.
Published: (2023)
Statistical Guarantees for High-Dimensional Stochastic Gradient Descent
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
Learning Tree-Based Models with Gradient Descent
by: Marton, Sascha
Published: (2026)
by: Marton, Sascha
Published: (2026)
Occam Gradient Descent
by: Kausik, B. N.
Published: (2024)
by: Kausik, B. N.
Published: (2024)
In-context Learning and Gradient Descent Revisited
by: Deutch, Gilad, et al.
Published: (2023)
by: Deutch, Gilad, et al.
Published: (2023)
Learning Associative Memories with Gradient Descent
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Efficient Search for Customized Activation Functions with Gradient Descent
by: Strack, Lukas, et al.
Published: (2024)
by: Strack, Lukas, et al.
Published: (2024)
Functional Central Limit Theorem for Stochastic Gradient Descent
by: Flamand, Kessang, et al.
Published: (2026)
by: Flamand, Kessang, et al.
Published: (2026)
Personalized Federated Learning with Exact Stochastic Gradient Descent
by: Nikoloutsopoulos, Sotirios, et al.
Published: (2022)
by: Nikoloutsopoulos, Sotirios, et al.
Published: (2022)
Towards Learning Stochastic Population Models by Gradient Descent
by: Kreikemeyer, Justin N., et al.
Published: (2024)
by: Kreikemeyer, Justin N., et al.
Published: (2024)
Gradient Descent with Provably Tuned Learning-rate Schedules
by: Sharma, Dravyansh
Published: (2025)
by: Sharma, Dravyansh
Published: (2025)
Learning Curves of Stochastic Gradient Descent in Kernel Regression
by: Zhang, Haihan, et al.
Published: (2025)
by: Zhang, Haihan, et al.
Published: (2025)
Partially Lazy Gradient Descent for Smoothed Online Learning
by: Mhaisen, Naram, et al.
Published: (2026)
by: Mhaisen, Naram, et al.
Published: (2026)
Hybrid Coordinate Descent for Efficient Neural Network Learning Using Line Search and Gradient Descent
by: Hsiao, Yen-Che, et al.
Published: (2024)
by: Hsiao, Yen-Che, et al.
Published: (2024)
Stacking as Accelerated Gradient Descent
by: Agarwal, Naman, et al.
Published: (2024)
by: Agarwal, Naman, et al.
Published: (2024)
Stochastic Adaptive Gradient Descent Without Descent
by: Aujol, Jean-François, et al.
Published: (2025)
by: Aujol, Jean-François, et al.
Published: (2025)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
Enhancing Fractional Gradient Descent with Learned Optimizers
by: Sobotka, Jan, et al.
Published: (2025)
by: Sobotka, Jan, et al.
Published: (2025)
Corner Gradient Descent
by: Yarotsky, Dmitry
Published: (2025)
by: Yarotsky, Dmitry
Published: (2025)
Continuum Transformers Perform In-Context Learning by Operator Gradient Descent
by: Mishra, Abhiti, et al.
Published: (2025)
by: Mishra, Abhiti, et al.
Published: (2025)
FocusLearn: Fully-Interpretable, High-Performance Modular Neural Networks for Time Series
by: Su, Qiqi, et al.
Published: (2023)
by: Su, Qiqi, et al.
Published: (2023)
Is All Learning (Natural) Gradient Descent?
by: Shoji, Lucas, et al.
Published: (2024)
by: Shoji, Lucas, et al.
Published: (2024)
Similar Items
-
Deep Linear Discriminant Analysis Revisited
by: Tezekbayev, Maxat, et al.
Published: (2026) -
Simplex Deep Linear Discriminant Analysis
by: Tezekbayev, Maxat, et al.
Published: (2026) -
Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel Thresholding
by: Takhanov, Rustem, et al.
Published: (2026) -
Overspecified Mixture Discriminant Analysis: Exponential Convergence, Statistical Guarantees, and Remote Sensing Applications
by: Bolatov, Arman, et al.
Published: (2025) -
Classifier Performance on Long‐Tail Distributions
by: Artur Pak, et al.
Published: (2025)