Unified theory of upper confidence bound policies for bandit problems targeting total reward, maximal reward, and more

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kikkawa, Nobuaki, Ohno, Hiroshi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!