Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ustimenko, Aleksei, Beznosikov, Aleksandr
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916185114673152
author Ustimenko, Aleksei
Beznosikov, Aleksandr
author_facet Ustimenko, Aleksei
Beznosikov, Aleksandr
contents In this work, we consider rather general and broad class of Markov chains, Ito chains, that look like Euler-Maryama discretization of some Stochastic Differential Equation. The chain we study is a unified framework for theoretical analysis. It comes with almost arbitrary isotropic and state-dependent noise instead of normal and state-independent one as in most related papers. Moreover, in our chain the drift and diffusion coefficient can be inexact in order to cover wide range of applications as Stochastic Gradient Langevin Dynamics, sampling, Stochastic Gradient Descent or Stochastic Gradient Boosting. We prove the bound in $W_{2}$-distance between the laws of our Ito chain and corresponding differential equation. These results improve or cover most of the known estimates. And for some particular cases, our analysis is the first.
format Preprint
id arxiv_https___arxiv_org_abs_2310_06081
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting
Ustimenko, Aleksei
Beznosikov, Aleksandr
Optimization and Control
Machine Learning
Probability
In this work, we consider rather general and broad class of Markov chains, Ito chains, that look like Euler-Maryama discretization of some Stochastic Differential Equation. The chain we study is a unified framework for theoretical analysis. It comes with almost arbitrary isotropic and state-dependent noise instead of normal and state-independent one as in most related papers. Moreover, in our chain the drift and diffusion coefficient can be inexact in order to cover wide range of applications as Stochastic Gradient Langevin Dynamics, sampling, Stochastic Gradient Descent or Stochastic Gradient Boosting. We prove the bound in $W_{2}$-distance between the laws of our Ito chain and corresponding differential equation. These results improve or cover most of the known estimates. And for some particular cases, our analysis is the first.
title Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting
topic Optimization and Control
Machine Learning
Probability
url https://arxiv.org/abs/2310.06081