Rectifying Regression in Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ayoub, Alex, Szepesvári, David, Bakhtiari, Alireza, Szepesvári, Csaba, Schuurmans, Dale
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917070207188992
author Ayoub, Alex
Szepesvári, David
Bakhtiari, Alireza
Szepesvári, Csaba
Schuurmans, Dale
author_facet Ayoub, Alex
Szepesvári, David
Bakhtiari, Alireza
Szepesvári, Csaba
Schuurmans, Dale
contents This paper investigates the impact of the loss function in value-based methods for reinforcement learning through an analysis of underlying prediction objectives. We theoretically show that mean absolute error is a better prediction objective than the traditional mean squared error for controlling the learned policy's suboptimality gap. Furthermore, we present results that different loss functions are better aligned with these different regression objectives: binary and categorical cross-entropy losses with the mean absolute error and squared loss with the mean squared error. We then provide empirical evidence that algorithms minimizing these cross-entropy losses can outperform those based on the squared loss in linear reinforcement learning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00885
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rectifying Regression in Reinforcement Learning
Ayoub, Alex
Szepesvári, David
Bakhtiari, Alireza
Szepesvári, Csaba
Schuurmans, Dale
Machine Learning
This paper investigates the impact of the loss function in value-based methods for reinforcement learning through an analysis of underlying prediction objectives. We theoretically show that mean absolute error is a better prediction objective than the traditional mean squared error for controlling the learned policy's suboptimality gap. Furthermore, we present results that different loss functions are better aligned with these different regression objectives: binary and categorical cross-entropy losses with the mean absolute error and squared loss with the mean squared error. We then provide empirical evidence that algorithms minimizing these cross-entropy losses can outperform those based on the squared loss in linear reinforcement learning.
title Rectifying Regression in Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2510.00885