A geometric framework for momentum-based optimizers for low-rank training
Fuente:
arXiv
Saved in:
| Main Authors: | Schotthöfer, Steffen, Klein, Timon, Kusch, Jonas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tucker Attention: A generalization of approximate attention mechanisms
by: Klein, Timon, et al.
Published: (2026)
by: Klein, Timon, et al.
Published: (2026)
An Augmented Backward-Corrected Projector Splitting Integrator for Dynamical Low-Rank Training
by: Kusch, Jonas, et al.
Published: (2025)
by: Kusch, Jonas, et al.
Published: (2025)
Geometry-aware training of factorized layers in tensor Tucker format
by: Zangrando, Emanuele, et al.
Published: (2023)
by: Zangrando, Emanuele, et al.
Published: (2023)
Mitigating Subject Dependency in EEG Decoding with Subject-Specific Low-Rank Adapters
by: Klein, Timon, et al.
Published: (2025)
by: Klein, Timon, et al.
Published: (2025)
Construction of high-order conservative basis-update and Galerkin dynamical low-rank integrators
by: Einkemmer, Lukas, et al.
Published: (2023)
by: Einkemmer, Lukas, et al.
Published: (2023)
GeoLoRA: Geometric integration for parameter efficient fine-tuning
by: Schotthöfer, Steffen, et al.
Published: (2024)
by: Schotthöfer, Steffen, et al.
Published: (2024)
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
by: Schotthöfer, Steffen, et al.
Published: (2024)
by: Schotthöfer, Steffen, et al.
Published: (2024)
Dynamical Low-Rank Compression of Neural Networks with Robustness under Adversarial Attacks
by: Schotthöfer, Steffen, et al.
Published: (2025)
by: Schotthöfer, Steffen, et al.
Published: (2025)
Structure-preserving neural networks for the regularized entropy-based closure of the Boltzmann moment system
by: Schotthöfer, Steffen, et al.
Published: (2024)
by: Schotthöfer, Steffen, et al.
Published: (2024)
A new perspective on low-rank optimization
by: Bertsimas, Dimitris, et al.
Published: (2021)
by: Bertsimas, Dimitris, et al.
Published: (2021)
Optimal low-rank stochastic gradient estimation for LLM training
by: Li, Zehao, et al.
Published: (2026)
by: Li, Zehao, et al.
Published: (2026)
Second-order robust parallel integrators for dynamical low-rank approximation
by: Kusch, Jonas
Published: (2024)
by: Kusch, Jonas
Published: (2024)
Worst-case low-rank approximations
by: Fries, Anya, et al.
Published: (2026)
by: Fries, Anya, et al.
Published: (2026)
Robust low-rank training via approximate orthonormal constraints
by: Savostianova, Dayana, et al.
Published: (2023)
by: Savostianova, Dayana, et al.
Published: (2023)
HADL Framework for Noise Resilient Long-Term Time Series Forecasting
by: Dey, Aditya, et al.
Published: (2025)
by: Dey, Aditya, et al.
Published: (2025)
Overshoot: Taking advantage of future gradients in momentum-based stochastic optimization
by: Kopal, Jakub, et al.
Published: (2025)
by: Kopal, Jakub, et al.
Published: (2025)
Geometrical structures of digital fluctuations in parameter space of neural networks trained with adaptive momentum optimization
by: Netay, Igor V.
Published: (2024)
by: Netay, Igor V.
Published: (2024)
Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers
by: Chertkov, Andrei, et al.
Published: (2025)
by: Chertkov, Andrei, et al.
Published: (2025)
A space-decoupling framework for optimization on bounded-rank matrices with orthogonally invariant constraints
by: Yang, Yan, et al.
Published: (2025)
by: Yang, Yan, et al.
Published: (2025)
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
A generalizable framework for low-rank tensor completion with numerical priors
by: Yuan, Shiran, et al.
Published: (2023)
by: Yuan, Shiran, et al.
Published: (2023)
An adaptive dynamical low-rank optimizer for solving kinetic parameter identification inverse problems
by: Baumann, Lena, et al.
Published: (2025)
by: Baumann, Lena, et al.
Published: (2025)
Accelerating nuclear-norm regularized low-rank matrix optimization through Burer-Monteiro decomposition
by: Lee, Ching-pei, et al.
Published: (2022)
by: Lee, Ching-pei, et al.
Published: (2022)
ARMAX identification of low rank graphical models
by: Cao, Wenqi, et al.
Published: (2025)
by: Cao, Wenqi, et al.
Published: (2025)
Batch, match, and patch: low-rank approximations for score-based variational inference
by: Modi, Chirag, et al.
Published: (2024)
by: Modi, Chirag, et al.
Published: (2024)
Weight decay induces low-rank attention layers
by: Kobayashi, Seijin, et al.
Published: (2024)
by: Kobayashi, Seijin, et al.
Published: (2024)
LOST: Low-rank and Sparse Pre-training for Large Language Models
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
A robust second-order low-rank BUG integrator based on the midpoint rule
by: Ceruti, Gianluca, et al.
Published: (2024)
by: Ceruti, Gianluca, et al.
Published: (2024)
Globally optimized SVD compression of LLMs via Fermi-function-based rank selection and gauge fixing
by: Rausch, Roman, et al.
Published: (2025)
by: Rausch, Roman, et al.
Published: (2025)
BOtied: Multi-objective Bayesian optimization with tied multivariate ranks
by: Park, Ji Won, et al.
Published: (2023)
by: Park, Ji Won, et al.
Published: (2023)
A simulation-based training framework for machine-learning applications in ARPES
by: Na, MengXing, et al.
Published: (2025)
by: Na, MengXing, et al.
Published: (2025)
Linear Recursive Feature Machines provably recover low-rank matrices
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
Convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Kolmogorov Arnold Informed neural network: A physics-informed deep learning framework for solving forward and inverse problems based on Kolmogorov Arnold Networks
by: Wang, Yizheng, et al.
Published: (2024)
by: Wang, Yizheng, et al.
Published: (2024)
Safety and optimality in learning-based control at low computational cost
by: Baumann, Dominik, et al.
Published: (2025)
by: Baumann, Dominik, et al.
Published: (2025)
Randomized Kaczmarz with geometrically smoothed momentum
by: Alderman, Seth J., et al.
Published: (2024)
by: Alderman, Seth J., et al.
Published: (2024)
A unified framework for hard and soft clustering with regularized optimal transport
by: Diebold, Jean-Frédéric, et al.
Published: (2017)
by: Diebold, Jean-Frédéric, et al.
Published: (2017)
A framework for measuring the training efficiency of a neural architecture
by: Cueto-Mendoza, Eduardo, et al.
Published: (2024)
by: Cueto-Mendoza, Eduardo, et al.
Published: (2024)
Sensor optimization for urban wind estimation with cluster-based probabilistic framework
by: Liang, Yutong, et al.
Published: (2025)
by: Liang, Yutong, et al.
Published: (2025)
Using Petri Nets as an Integrated Constraint Mechanism for Reinforcement Learning Tasks
by: Sachweh, Timon, et al.
Published: (2024)
by: Sachweh, Timon, et al.
Published: (2024)
Similar Items
-
Tucker Attention: A generalization of approximate attention mechanisms
by: Klein, Timon, et al.
Published: (2026) -
An Augmented Backward-Corrected Projector Splitting Integrator for Dynamical Low-Rank Training
by: Kusch, Jonas, et al.
Published: (2025) -
Geometry-aware training of factorized layers in tensor Tucker format
by: Zangrando, Emanuele, et al.
Published: (2023) -
Mitigating Subject Dependency in EEG Decoding with Subject-Specific Low-Rank Adapters
by: Klein, Timon, et al.
Published: (2025) -
Construction of high-order conservative basis-update and Galerkin dynamical low-rank integrators
by: Einkemmer, Lukas, et al.
Published: (2023)