Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Zhenzhang, Peyré, Gabriel, Cremers, Daniel, Ablin, Pierre |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Smooth Is Attention?
by: Castin, Valérie, et al.
Published: (2023)
by: Castin, Valérie, et al.
Published: (2023)
Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence
by: Castin, Valérie, et al.
Published: (2026)
by: Castin, Valérie, et al.
Published: (2026)
A Unified Perspective on the Dynamics of Deep Transformers
by: Castin, Valérie, et al.
Published: (2025)
by: Castin, Valérie, et al.
Published: (2025)
Glocal Hypergradient Estimation with Koopman Operator
by: Hataya, Ryuichiro, et al.
Published: (2024)
by: Hataya, Ryuichiro, et al.
Published: (2024)
Robust Sublinear Convergence Rates for Iterative Bregman Projections
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
Convergence Properties of Stochastic Hypergradients
by: Grazzi, Riccardo, et al.
Published: (2020)
by: Grazzi, Riccardo, et al.
Published: (2020)
Muon Dynamics as a Spectral Wasserstein Flow
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
Optimal and Diffusion Transports in Machine Learning
by: Peyré, Gabriel
Published: (2025)
by: Peyré, Gabriel
Published: (2025)
Optimal Transport for Machine Learners
by: Peyré, Gabriel
Published: (2025)
by: Peyré, Gabriel
Published: (2025)
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026)
by: Monteiro, João, et al.
Published: (2026)
Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent
by: Chu, Ya-Chi, et al.
Published: (2025)
by: Chu, Ya-Chi, et al.
Published: (2025)
Efficient Bilevel Optimization with KFAC-Based Hypergradients
by: Liao, Disen, et al.
Published: (2026)
by: Liao, Disen, et al.
Published: (2026)
The AdEMAMix Optimizer: Better, Faster, Older
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
Towards Understanding the Universality of Transformers for Next-Token Prediction
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
Intrinsic training dynamics of deep neural networks
by: Marcotte, Sibylle, et al.
Published: (2025)
by: Marcotte, Sibylle, et al.
Published: (2025)
Geometry-Aware Discretization Error of Diffusion Models
by: Hurault, Samuel, et al.
Published: (2026)
by: Hurault, Samuel, et al.
Published: (2026)
Transformative or Conservative? Conservation laws for ResNets and Transformers
by: Marcotte, Sibylle, et al.
Published: (2025)
by: Marcotte, Sibylle, et al.
Published: (2025)
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Efficient Curvature-Aware Hypergradient Approximation for Bilevel Optimization
by: Dong, Youran, et al.
Published: (2025)
by: Dong, Youran, et al.
Published: (2025)
Playing Markov Games Without Observing Payoffs
by: Ablin, Daniel, et al.
Published: (2025)
by: Ablin, Daniel, et al.
Published: (2025)
HYDRA: Hypergradient Data Relevance Analysis for Interpreting Deep Neural Networks
by: Chen, Yuanyuan, et al.
Published: (2021)
by: Chen, Yuanyuan, et al.
Published: (2021)
Locking Pretrained Weights via Deep Low-Rank Residual Distillation
by: Sakamoto, Keitaro, et al.
Published: (2026)
by: Sakamoto, Keitaro, et al.
Published: (2026)
MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations
by: Heurtebise, Ambroise, et al.
Published: (2025)
by: Heurtebise, Ambroise, et al.
Published: (2025)
Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
by: Kudo, Mikoto, et al.
Published: (2026)
by: Kudo, Mikoto, et al.
Published: (2026)
Keep the Momentum: Conservation Laws beyond Euclidean Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2024)
by: Marcotte, Sibylle, et al.
Published: (2024)
Learning from Samples: Inverse Problems over measures via Sharpened Fenchel-Young Losses
by: Andrade, Francisco, et al.
Published: (2025)
by: Andrade, Francisco, et al.
Published: (2025)
Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2023)
by: Marcotte, Sibylle, et al.
Published: (2023)
On the global convergence of gradient descent for wide shallow models with bounded nonlinearities
by: Petit, Romain, et al.
Published: (2026)
by: Petit, Romain, et al.
Published: (2026)
Federated Learning with Hypergradient-based Online Update of Aggregation Weights
by: Nakai-Kasai, Ayano, et al.
Published: (2026)
by: Nakai-Kasai, Ayano, et al.
Published: (2026)
A framework for bilevel optimization that enables stochastic and global variance reduction algorithms
by: Dagréou, Mathieu, et al.
Published: (2022)
by: Dagréou, Mathieu, et al.
Published: (2022)
A Lower Bound and a Near-Optimal Algorithm for Bilevel Empirical Risk Minimization
by: Dagréou, Mathieu, et al.
Published: (2023)
by: Dagréou, Mathieu, et al.
Published: (2023)
MedMamba: Vision Mamba for Medical Image Classification
by: Yue, Yubiao, et al.
Published: (2024)
by: Yue, Yubiao, et al.
Published: (2024)
Bi-Level Policy Optimization with Nyström Hypergradients
by: Prakash, Arjun, et al.
Published: (2025)
by: Prakash, Arjun, et al.
Published: (2025)
Natural Hypergradient Descent: Algorithm Design, Convergence Analysis, and Parallel Implementation
by: Kong, Deyi, et al.
Published: (2026)
by: Kong, Deyi, et al.
Published: (2026)
Preconditioned Attention: Enhancing Efficiency in Transformers
by: Saratchandran, Hemanth
Published: (2026)
by: Saratchandran, Hemanth
Published: (2026)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Need a Small Specialized Language Model? Plan Early!
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Scaling Laws for Mixture Pretraining Under Data Constraints
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)
by: Ablin, Pierre, et al.
Published: (2025)
Understanding the training of infinitely deep and wide ResNets with Conditional Optimal Transport
by: Barboni, Raphaël, et al.
Published: (2024)
by: Barboni, Raphaël, et al.
Published: (2024)
Similar Items
-
How Smooth Is Attention?
by: Castin, Valérie, et al.
Published: (2023) -
Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence
by: Castin, Valérie, et al.
Published: (2026) -
A Unified Perspective on the Dynamics of Deep Transformers
by: Castin, Valérie, et al.
Published: (2025) -
Glocal Hypergradient Estimation with Koopman Operator
by: Hataya, Ryuichiro, et al.
Published: (2024) -
Robust Sublinear Convergence Rates for Iterative Bregman Projections
by: Peyré, Gabriel
Published: (2026)