Large-Scale Riemannian Meta-Optimization via Subspace Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Peilin, Wu, Yuwei, Gao, Zhi, Fan, Xiaomeng, Jia, Yunde
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909479663042560
author Yu, Peilin
Wu, Yuwei
Gao, Zhi
Fan, Xiaomeng
Jia, Yunde
author_facet Yu, Peilin
Wu, Yuwei
Gao, Zhi
Fan, Xiaomeng
Jia, Yunde
contents Riemannian meta-optimization provides a promising approach to solving non-linear constrained optimization problems, which trains neural networks as optimizers to perform optimization on Riemannian manifolds. However, existing Riemannian meta-optimization methods take up huge memory footprints in large-scale optimization settings, as the learned optimizer can only adapt gradients of a fixed size and thus cannot be shared across different Riemannian parameters. In this paper, we propose an efficient Riemannian meta-optimization method that significantly reduces the memory burden for large-scale optimization via a subspace adaptation scheme. Our method trains neural networks to individually adapt the row and column subspaces of Riemannian gradients, instead of directly adapting the full gradient matrices in existing Riemannian meta-optimization methods. In this case, our learned optimizer can be shared across Riemannian parameters with different sizes. Our method reduces the model memory consumption by six orders of magnitude when optimizing an orthogonal mainstream deep neural network (e.g., ResNet50). Experiments on multiple Riemannian tasks show that our method can not only reduce the memory consumption but also improve the performance of Riemannian meta-optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15235
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large-Scale Riemannian Meta-Optimization via Subspace Adaptation
Yu, Peilin
Wu, Yuwei
Gao, Zhi
Fan, Xiaomeng
Jia, Yunde
Machine Learning
Computer Vision and Pattern Recognition
Riemannian meta-optimization provides a promising approach to solving non-linear constrained optimization problems, which trains neural networks as optimizers to perform optimization on Riemannian manifolds. However, existing Riemannian meta-optimization methods take up huge memory footprints in large-scale optimization settings, as the learned optimizer can only adapt gradients of a fixed size and thus cannot be shared across different Riemannian parameters. In this paper, we propose an efficient Riemannian meta-optimization method that significantly reduces the memory burden for large-scale optimization via a subspace adaptation scheme. Our method trains neural networks to individually adapt the row and column subspaces of Riemannian gradients, instead of directly adapting the full gradient matrices in existing Riemannian meta-optimization methods. In this case, our learned optimizer can be shared across Riemannian parameters with different sizes. Our method reduces the model memory consumption by six orders of magnitude when optimizing an orthogonal mainstream deep neural network (e.g., ResNet50). Experiments on multiple Riemannian tasks show that our method can not only reduce the memory consumption but also improve the performance of Riemannian meta-optimization.
title Large-Scale Riemannian Meta-Optimization via Subspace Adaptation
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.15235