DROGO: Default Representation Objective via Graph Optimization in Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tse, Hon Tik, Machado, Marlos C.
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914296554848256
author Tse, Hon Tik
Machado, Marlos C.
author_facet Tse, Hon Tik
Machado, Marlos C.
contents In computational reinforcement learning, the default representation (DR) and its principal eigenvector have been shown to be effective for a wide variety of applications, including reward shaping, count-based exploration, option discovery, and transfer. However, in prior investigations, the eigenvectors of the DR were computed by first approximating the DR matrix, and then performing an eigendecomposition. This procedure is computationally expensive and does not scale to high-dimensional spaces. In this paper, we derive an objective for directly approximating the principal eigenvector of the DR with a neural network. We empirically demonstrate the effectiveness of the objective in a number of environments, and apply the learned eigenvectors for reward shaping.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00403
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DROGO: Default Representation Objective via Graph Optimization in Reinforcement Learning
Tse, Hon Tik
Machado, Marlos C.
Machine Learning
In computational reinforcement learning, the default representation (DR) and its principal eigenvector have been shown to be effective for a wide variety of applications, including reward shaping, count-based exploration, option discovery, and transfer. However, in prior investigations, the eigenvectors of the DR were computed by first approximating the DR matrix, and then performing an eigendecomposition. This procedure is computationally expensive and does not scale to high-dimensional spaces. In this paper, we derive an objective for directly approximating the principal eigenvector of the DR with a neural network. We empirically demonstrate the effectiveness of the objective in a number of environments, and apply the learned eigenvectors for reward shaping.
title DROGO: Default Representation Objective via Graph Optimization in Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2602.00403