Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ayoub, Alex, Wang, Kaiwen, Liu, Vincent, Robertson, Samuel, McInerney, James, Liang, Dawen, Kallus, Nathan, Szepesvári, Csaba
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911974321815552
author Ayoub, Alex
Wang, Kaiwen
Liu, Vincent
Robertson, Samuel
McInerney, James
Liang, Dawen
Kallus, Nathan
Szepesvári, Csaba
author_facet Ayoub, Alex
Wang, Kaiwen
Liu, Vincent
Robertson, Samuel
McInerney, James
Liang, Dawen
Kallus, Nathan
Szepesvári, Csaba
contents We propose training fitted Q-iteration with log-loss (FQI-log) for batch reinforcement learning (RL). We show that the number of samples needed to learn a near-optimal policy with FQI-log scales with the accumulated cost of the optimal policy, which is zero in problems where acting optimally achieves the goal and incurs no cost. In doing so, we provide a general framework for proving small-cost bounds, i.e. bounds that scale with the optimal achievable cost, in batch RL. Moreover, we empirically verify that FQI-log uses fewer samples than FQI trained with squared loss on problems where the optimal policy reliably achieves the goal.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05385
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
Ayoub, Alex
Wang, Kaiwen
Liu, Vincent
Robertson, Samuel
McInerney, James
Liang, Dawen
Kallus, Nathan
Szepesvári, Csaba
Machine Learning
We propose training fitted Q-iteration with log-loss (FQI-log) for batch reinforcement learning (RL). We show that the number of samples needed to learn a near-optimal policy with FQI-log scales with the accumulated cost of the optimal policy, which is zero in problems where acting optimally achieves the goal and incurs no cost. In doing so, we provide a general framework for proving small-cost bounds, i.e. bounds that scale with the optimal achievable cost, in batch RL. Moreover, we empirically verify that FQI-log uses fewer samples than FQI trained with squared loss on problems where the optimal policy reliably achieves the goal.
title Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2403.05385