Evaluation-Aware Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deshmukh, Shripad Vilasrao, Schwarzer, Will, Niekum, Scott
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908902459703296
author Deshmukh, Shripad Vilasrao
Schwarzer, Will
Niekum, Scott
author_facet Deshmukh, Shripad Vilasrao
Schwarzer, Will
Niekum, Scott
contents Policy evaluation is a core component of many reinforcement learning (RL) algorithms and a critical tool for ensuring safe deployment of RL policies. However, existing policy evaluation methods often suffer from high variance or bias. To address these issues, we introduce Evaluation-Aware Reinforcement Learning (EvA-RL), a general policy learning framework that considers evaluation accuracy at train-time, as opposed to standard post-hoc policy evaluation methods. Specifically, EvA-RL directly optimizes policies for efficient and accurate evaluation, in addition to being performant. We provide an instantiation of EvA-RL and demonstrate through a combination of theoretical analysis and empirical results that EvA-RL effectively trades off between evaluation accuracy and expected return. Finally, we show that the evaluation-aware policy and the evaluation mechanism itself can be co-learned to mitigate this tradeoff, providing the evaluation benefits without significantly sacrificing policy performance. This work opens a new line of research that elevates reliable evaluation to a first-class principle in reinforcement learning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19464
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluation-Aware Reinforcement Learning
Deshmukh, Shripad Vilasrao
Schwarzer, Will
Niekum, Scott
Artificial Intelligence
Machine Learning
Policy evaluation is a core component of many reinforcement learning (RL) algorithms and a critical tool for ensuring safe deployment of RL policies. However, existing policy evaluation methods often suffer from high variance or bias. To address these issues, we introduce Evaluation-Aware Reinforcement Learning (EvA-RL), a general policy learning framework that considers evaluation accuracy at train-time, as opposed to standard post-hoc policy evaluation methods. Specifically, EvA-RL directly optimizes policies for efficient and accurate evaluation, in addition to being performant. We provide an instantiation of EvA-RL and demonstrate through a combination of theoretical analysis and empirical results that EvA-RL effectively trades off between evaluation accuracy and expected return. Finally, we show that the evaluation-aware policy and the evaluation mechanism itself can be co-learned to mitigate this tradeoff, providing the evaluation benefits without significantly sacrificing policy performance. This work opens a new line of research that elevates reliable evaluation to a first-class principle in reinforcement learning.
title Evaluation-Aware Reinforcement Learning
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.19464