Improved Stochastic Optimization of LogSumExp

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gladin, Egor, Kroshnin, Alexey, Zhu, Jia-Jie, Dvurechensky, Pavel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908808980201472
author Gladin, Egor
Kroshnin, Alexey
Zhu, Jia-Jie
Dvurechensky, Pavel
author_facet Gladin, Egor
Kroshnin, Alexey
Zhu, Jia-Jie
Dvurechensky, Pavel
contents The LogSumExp function, dual to the Kullback-Leibler (KL) divergence, plays a central role in many important optimization problems, including entropy-regularized optimal transport (OT) and distributionally robust optimization (DRO). In practice, when the number of exponential terms inside the logarithm is large or infinite, optimization becomes challenging since computing the gradient requires differentiating every term. We propose a novel convexity- and smoothness-preserving approximation to LogSumExp that can be efficiently optimized using stochastic gradient methods. This approximation is rooted in a sound modification of the KL divergence in the dual, resulting in a new $f$-divergence called the safe KL divergence. Our experiments and theoretical analysis of the LogSumExp-based stochastic optimization, arising in DRO and continuous OT, demonstrate the advantages of our approach over existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24894
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improved Stochastic Optimization of LogSumExp
Gladin, Egor
Kroshnin, Alexey
Zhu, Jia-Jie
Dvurechensky, Pavel
Optimization and Control
Machine Learning
The LogSumExp function, dual to the Kullback-Leibler (KL) divergence, plays a central role in many important optimization problems, including entropy-regularized optimal transport (OT) and distributionally robust optimization (DRO). In practice, when the number of exponential terms inside the logarithm is large or infinite, optimization becomes challenging since computing the gradient requires differentiating every term. We propose a novel convexity- and smoothness-preserving approximation to LogSumExp that can be efficiently optimized using stochastic gradient methods. This approximation is rooted in a sound modification of the KL divergence in the dual, resulting in a new $f$-divergence called the safe KL divergence. Our experiments and theoretical analysis of the LogSumExp-based stochastic optimization, arising in DRO and continuous OT, demonstrate the advantages of our approach over existing baselines.
title Improved Stochastic Optimization of LogSumExp
topic Optimization and Control
Machine Learning
url https://arxiv.org/abs/2509.24894