A Case for Validation Buffer in Pessimistic Actor-Critic

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nauman, Michal, Ostaszewski, Mateusz, Cygan, Marek
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909125864062976
author Nauman, Michal
Ostaszewski, Mateusz
Cygan, Marek
author_facet Nauman, Michal
Ostaszewski, Mateusz
Cygan, Marek
contents In this paper, we investigate the issue of error accumulation in critic networks updated via pessimistic temporal difference objectives. We show that the critic approximation error can be approximated via a recursive fixed-point model similar to that of the Bellman value. We use such recursive definition to retrieve the conditions under which the pessimistic critic is unbiased. Building on these insights, we propose Validation Pessimism Learning (VPL) algorithm. VPL uses a small validation buffer to adjust the levels of pessimism throughout the agent training, with the pessimism set such that the approximation error of the critic targets is minimized. We investigate the proposed approach on a variety of locomotion and manipulation tasks and report improvements in sample efficiency and performance.
format Preprint
id arxiv_https___arxiv_org_abs_2403_01014
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Case for Validation Buffer in Pessimistic Actor-Critic
Nauman, Michal
Ostaszewski, Mateusz
Cygan, Marek
Machine Learning
In this paper, we investigate the issue of error accumulation in critic networks updated via pessimistic temporal difference objectives. We show that the critic approximation error can be approximated via a recursive fixed-point model similar to that of the Bellman value. We use such recursive definition to retrieve the conditions under which the pessimistic critic is unbiased. Building on these insights, we propose Validation Pessimism Learning (VPL) algorithm. VPL uses a small validation buffer to adjust the levels of pessimism throughout the agent training, with the pessimism set such that the approximation error of the critic targets is minimized. We investigate the proposed approach on a variety of locomotion and manipulation tasks and report improvements in sample efficiency and performance.
title A Case for Validation Buffer in Pessimistic Actor-Critic
topic Machine Learning
url https://arxiv.org/abs/2403.01014