Dynamic Low-rank Approximation of Full-Matrix Preconditioner for Training Generalized Linear Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Matveeva, Tatyana, Katrutsa, Aleksandr, Frolov, Evgeny
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909758457380864
author Matveeva, Tatyana
Katrutsa, Aleksandr
Frolov, Evgeny
author_facet Matveeva, Tatyana
Katrutsa, Aleksandr
Frolov, Evgeny
contents Adaptive gradient methods like Adagrad and its variants are widespread in large-scale optimization. However, their use of diagonal preconditioning matrices limits the ability to capture parameter correlations. Full-matrix adaptive methods, approximating the exact Hessian, can model these correlations and may enable faster convergence. At the same time, their computational and memory costs are often prohibitive for large-scale models. To address this limitation, we propose AdaGram, an optimizer that enables efficient full-matrix adaptive gradient updates. To reduce memory and computational overhead, we utilize fast symmetric factorization for computing the preconditioned update direction at each iteration. Additionally, we maintain the low-rank structure of a preconditioner along the optimization trajectory using matrix integrator methods. Numerical experiments on standard machine learning tasks show that AdaGram converges faster or matches the performance of diagonal adaptive optimizers when using rank five and smaller rank approximations. This demonstrates AdaGram's potential as a scalable solution for adaptive optimization in large models.
format Preprint
id arxiv_https___arxiv_org_abs_2508_21106
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Low-rank Approximation of Full-Matrix Preconditioner for Training Generalized Linear Models
Matveeva, Tatyana
Katrutsa, Aleksandr
Frolov, Evgeny
Machine Learning
Artificial Intelligence
Adaptive gradient methods like Adagrad and its variants are widespread in large-scale optimization. However, their use of diagonal preconditioning matrices limits the ability to capture parameter correlations. Full-matrix adaptive methods, approximating the exact Hessian, can model these correlations and may enable faster convergence. At the same time, their computational and memory costs are often prohibitive for large-scale models. To address this limitation, we propose AdaGram, an optimizer that enables efficient full-matrix adaptive gradient updates. To reduce memory and computational overhead, we utilize fast symmetric factorization for computing the preconditioned update direction at each iteration. Additionally, we maintain the low-rank structure of a preconditioner along the optimization trajectory using matrix integrator methods. Numerical experiments on standard machine learning tasks show that AdaGram converges faster or matches the performance of diagonal adaptive optimizers when using rank five and smaller rank approximations. This demonstrates AdaGram's potential as a scalable solution for adaptive optimization in large models.
title Dynamic Low-rank Approximation of Full-Matrix Preconditioner for Training Generalized Linear Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.21106