A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Lattimore, Tor
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912985239257088
author Lattimore, Tor
author_facet Lattimore, Tor
contents We adapt the analysis of policy gradient for continuous time $k$-armed stochastic bandits by Lattimore (2026) to the standard discrete time setup. As in continuous time, we prove that with learning rate $η= O(Δ_{\min}^2/(Δ_{\max} \log(n)))$ the regret is $O(k \log(k) \log(n) / η)$ where $n$ is the horizon and $Δ_{\min}$ and $Δ_{\max}$ are the minimum and maximum gaps.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26547
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
Lattimore, Tor
Machine Learning
We adapt the analysis of policy gradient for continuous time $k$-armed stochastic bandits by Lattimore (2026) to the standard discrete time setup. As in continuous time, we prove that with learning rate $η= O(Δ_{\min}^2/(Δ_{\max} \log(n)))$ the regret is $O(k \log(k) \log(n) / η)$ where $n$ is the horizon and $Δ_{\min}$ and $Δ_{\max}$ are the minimum and maximum gaps.
title A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
topic Machine Learning
url https://arxiv.org/abs/2603.26547