Layerwise LQR for Geometry-Aware Optimization of Deep Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dufort-Labbé, Simon, Bacon, Pierre-Luc, Pascanu, Razvan, Lacoste-Julien, Simon, Baratin, Aristide
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913092014702592
author Dufort-Labbé, Simon
Bacon, Pierre-Luc
Pascanu, Razvan
Lacoste-Julien, Simon
Baratin, Aristide
author_facet Dufort-Labbé, Simon
Bacon, Pierre-Luc
Pascanu, Razvan
Lacoste-Julien, Simon
Baratin, Aristide
contents Geometry-aware optimizers such as Newton and natural gradient can improve conditioning in deep learning, but scalable variants such as K-FAC, Shampoo, and related preconditioners usually impose structural approximations early, often discarding cross-layer interactions induced by the network computation. We introduce Layerwise LQR (LLQR), a framework for learning structured inverse preconditioners under a global layerwise optimal-control objective. The starting point is an exact equivalence: the steepest-descent step under a broad class of divergence-induced quadratic models--including Newton, Gauss-Newton, Fisher/natural-gradient, and intermediate-layer metrics--can be written as a finite-horizon Linear Quadratic Regulator (LQR) problem. This formulation serves as a reference that exposes the layerwise dynamics and cost matrices encoding the original dense geometry. We then derive a scalable relaxation that learns diagonal, (E-)Kronecker-factored, or other structured inverse preconditioners by minimizing the LQR objective and reusing them across iterations. The resulting optimizer wraps standard methods while retaining a principled connection to second-order geometry, without forming or inverting the global curvature matrix. Experiments on ResNets and Transformers show that LLQR improves optimization dynamics and often translates these gains into improved final test performance, while adding only modest wall-clock overhead. It establishes LLQR as a practical framework for geometry-aware second-order methods and a reference for evaluating scalable approximations.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04230
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Layerwise LQR for Geometry-Aware Optimization of Deep Networks
Dufort-Labbé, Simon
Bacon, Pierre-Luc
Pascanu, Razvan
Lacoste-Julien, Simon
Baratin, Aristide
Machine Learning
Artificial Intelligence
Geometry-aware optimizers such as Newton and natural gradient can improve conditioning in deep learning, but scalable variants such as K-FAC, Shampoo, and related preconditioners usually impose structural approximations early, often discarding cross-layer interactions induced by the network computation. We introduce Layerwise LQR (LLQR), a framework for learning structured inverse preconditioners under a global layerwise optimal-control objective. The starting point is an exact equivalence: the steepest-descent step under a broad class of divergence-induced quadratic models--including Newton, Gauss-Newton, Fisher/natural-gradient, and intermediate-layer metrics--can be written as a finite-horizon Linear Quadratic Regulator (LQR) problem. This formulation serves as a reference that exposes the layerwise dynamics and cost matrices encoding the original dense geometry. We then derive a scalable relaxation that learns diagonal, (E-)Kronecker-factored, or other structured inverse preconditioners by minimizing the LQR objective and reusing them across iterations. The resulting optimizer wraps standard methods while retaining a principled connection to second-order geometry, without forming or inverting the global curvature matrix. Experiments on ResNets and Transformers show that LLQR improves optimization dynamics and often translates these gains into improved final test performance, while adding only modest wall-clock overhead. It establishes LLQR as a practical framework for geometry-aware second-order methods and a reference for evaluating scalable approximations.
title Layerwise LQR for Geometry-Aware Optimization of Deep Networks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.04230