Saved in:
Bibliographic Details
Main Authors: Gan, Min, Chen, Guang-Yong, Yi, Yang, Yang, Lin
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.01234
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914131056001024
author Gan, Min
Chen, Guang-Yong
Yi, Yang
Yang, Lin
author_facet Gan, Min
Chen, Guang-Yong
Yi, Yang
Yang, Lin
contents The proliferation of saddle points, rather than poor local minima, is increasingly understood to be a primary obstacle in large-scale non-convex optimization for machine learning. Variable elimination algorithms, like Variable Projection (VarPro), have long been observed to exhibit superior convergence and robustness in practice, yet a principled understanding of why they so effectively navigate these complex energy landscapes has remained elusive. In this work, we provide a rigorous geometric explanation by comparing the optimization landscapes of the original and reduced formulations. Through a rigorous analysis based on Hessian inertia and the Schur complement, we prove that variable elimination fundamentally reshapes the critical point structure of the objective function, revealing that local maxima in the reduced landscape are created from, and correspond directly to, saddle points in the original formulation. Our findings are illustrated on the canonical problem of non-convex matrix factorization, visualized directly on two-parameter neural networks, and finally validated in training deep Residual Networks, where our approach yields dramatic improvements in stability and convergence to superior minima. This work goes beyond explaining an existing method; it establishes landscape simplification via saddle point transformation as a powerful principle that can guide the design of a new generation of more robust and efficient optimization algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01234
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Saddle Point Remedy: Power of Variable Elimination in Non-convex Optimization
Gan, Min
Chen, Guang-Yong
Yi, Yang
Yang, Lin
Machine Learning
The proliferation of saddle points, rather than poor local minima, is increasingly understood to be a primary obstacle in large-scale non-convex optimization for machine learning. Variable elimination algorithms, like Variable Projection (VarPro), have long been observed to exhibit superior convergence and robustness in practice, yet a principled understanding of why they so effectively navigate these complex energy landscapes has remained elusive. In this work, we provide a rigorous geometric explanation by comparing the optimization landscapes of the original and reduced formulations. Through a rigorous analysis based on Hessian inertia and the Schur complement, we prove that variable elimination fundamentally reshapes the critical point structure of the objective function, revealing that local maxima in the reduced landscape are created from, and correspond directly to, saddle points in the original formulation. Our findings are illustrated on the canonical problem of non-convex matrix factorization, visualized directly on two-parameter neural networks, and finally validated in training deep Residual Networks, where our approach yields dramatic improvements in stability and convergence to superior minima. This work goes beyond explaining an existing method; it establishes landscape simplification via saddle point transformation as a powerful principle that can guide the design of a new generation of more robust and efficient optimization algorithms.
title A Saddle Point Remedy: Power of Variable Elimination in Non-convex Optimization
topic Machine Learning
url https://arxiv.org/abs/2511.01234