The informativeness of the gradient revisited
Fuente:
arXiv
Saved in:
| Main Author: | Takhanov, Rustem |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Intrinsic Dimensions of Data in Kernel Learning
by: Takhanov, Rustem
Published: (2026)
by: Takhanov, Rustem
Published: (2026)
Multi-layer random features and the approximation power of neural networks
by: Takhanov, Rustem
Published: (2024)
by: Takhanov, Rustem
Published: (2024)
Non-asymptotic spectral bounds on the $\varepsilon$-entropy of kernel classes
by: Takhanov, Rustem
Published: (2022)
by: Takhanov, Rustem
Published: (2022)
Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel Thresholding
by: Takhanov, Rustem, et al.
Published: (2026)
by: Takhanov, Rustem, et al.
Published: (2026)
Deep Linear Discriminant Analysis Revisited
by: Tezekbayev, Maxat, et al.
Published: (2026)
by: Tezekbayev, Maxat, et al.
Published: (2026)
Gradient Descent Fails to Learn High-frequency Functions and Modular Arithmetic
by: Takhanov, Rustem, et al.
Published: (2023)
by: Takhanov, Rustem, et al.
Published: (2023)
Safe-EF: Error Feedback for Nonsmooth Constrained Optimization
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
Why Do We Need Warm-up? A Theoretical Perspective
by: Alimisis, Foivos, et al.
Published: (2025)
by: Alimisis, Foivos, et al.
Published: (2025)
Unbiased and Sign Compression in Distributed Learning: Comparing Noise Resilience via SDEs
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
Optimal ridge regularization revisited
by: Timmermans, Jack, et al.
Published: (2026)
by: Timmermans, Jack, et al.
Published: (2026)
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
by: Islamov, Rustem, et al.
Published: (2024)
by: Islamov, Rustem, et al.
Published: (2024)
Towards Faster Decentralized Stochastic Optimization with Communication Compression
by: Islamov, Rustem, et al.
Published: (2024)
by: Islamov, Rustem, et al.
Published: (2024)
Double Momentum and Error Feedback for Clipping with Fast Rates and Differential Privacy
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
Constants of motion network revisited
by: Fang, Wenqi, et al.
Published: (2025)
by: Fang, Wenqi, et al.
Published: (2025)
On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
Adaptive Methods Are Preferable in High Privacy Settings: An SDE Perspective
by: Compagnoni, Enea Monzio, et al.
Published: (2026)
by: Compagnoni, Enea Monzio, et al.
Published: (2026)
Variance-sensitive Thompson sampling for generalised linear bandits, revisited
by: Perneczky, Tom, et al.
Published: (2026)
by: Perneczky, Tom, et al.
Published: (2026)
Feature selection revisited in the single-cell era
by: Yang, Pengyi, et al.
Published: (2021)
by: Yang, Pengyi, et al.
Published: (2021)
FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information
by: Hwang, Dongseong
Published: (2024)
by: Hwang, Dongseong
Published: (2024)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
On the under-reaching phenomenon in message-passing neural PDE solvers: revisiting the CFL condition
by: Tesan, Lucas, et al.
Published: (2025)
by: Tesan, Lucas, et al.
Published: (2025)
Equivalence of stochastic and deterministic policy gradients
by: Todorov, Emo
Published: (2025)
by: Todorov, Emo
Published: (2025)
Policy gradient methods for ordinal policies
by: Weinberger, Simón, et al.
Published: (2025)
by: Weinberger, Simón, et al.
Published: (2025)
Distributional value gradients for stochastic environments
by: Debes, Baptiste, et al.
Published: (2026)
by: Debes, Baptiste, et al.
Published: (2026)
Fast training of accurate physics-informed neural networks without gradient descent
by: Datar, Chinmay, et al.
Published: (2024)
by: Datar, Chinmay, et al.
Published: (2024)
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Bayesian policy gradient and actor-critic algorithms
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
Stochastic approximation in non-markovian environments revisited
by: Borkar, Vivek Shripad
Published: (2026)
by: Borkar, Vivek Shripad
Published: (2026)
Almost sure convergence rates of stochastic gradient methods under gradient domination
by: Weissmann, Simon, et al.
Published: (2024)
by: Weissmann, Simon, et al.
Published: (2024)
Improved physics-informed neural network in mitigating gradient related failures
by: Niu, Pancheng, et al.
Published: (2024)
by: Niu, Pancheng, et al.
Published: (2024)
Semi-gradient DICE for Offline Constrained Reinforcement Learning
by: Kim, Woosung, et al.
Published: (2025)
by: Kim, Woosung, et al.
Published: (2025)
Some remarks on gradient dominance and LQR policy optimization
by: Sontag, Eduardo D.
Published: (2025)
by: Sontag, Eduardo D.
Published: (2025)
ISOPO: Proximal policy gradients without pi-old
by: Abrahamsen, Nilin
Published: (2025)
by: Abrahamsen, Nilin
Published: (2025)
A projection-based framework for gradient-free and parallel learning
by: Bergmeister, Andreas, et al.
Published: (2025)
by: Bergmeister, Andreas, et al.
Published: (2025)
On a few pitfalls in KL divergence gradient estimation for RL
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Curly Flow Matching for Learning Non-gradient Field Dynamics
by: Petrović, Katarina, et al.
Published: (2025)
by: Petrović, Katarina, et al.
Published: (2025)
A policy gradient approach for optimization of smooth risk measures
by: Vijayan, Nithia, et al.
Published: (2022)
by: Vijayan, Nithia, et al.
Published: (2022)
Similar Items
-
On the Intrinsic Dimensions of Data in Kernel Learning
by: Takhanov, Rustem
Published: (2026) -
Multi-layer random features and the approximation power of neural networks
by: Takhanov, Rustem
Published: (2024) -
Non-asymptotic spectral bounds on the $\varepsilon$-entropy of kernel classes
by: Takhanov, Rustem
Published: (2022) -
Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel Thresholding
by: Takhanov, Rustem, et al.
Published: (2026) -
Deep Linear Discriminant Analysis Revisited
by: Tezekbayev, Maxat, et al.
Published: (2026)