Universal Value-Function Uncertainties

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zanger, Moritz A., Weltevrede, Max, Oren, Yaniv, Van der Vaart, Pascal R., Horsch, Caroline, Böhmer, Wendelin, Spaan, Matthijs T. J.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909631526207488
author Zanger, Moritz A.
Weltevrede, Max
Oren, Yaniv
Van der Vaart, Pascal R.
Horsch, Caroline
Böhmer, Wendelin
Spaan, Matthijs T. J.
author_facet Zanger, Moritz A.
Weltevrede, Max
Oren, Yaniv
Van der Vaart, Pascal R.
Horsch, Caroline
Böhmer, Wendelin
Spaan, Matthijs T. J.
contents Estimating epistemic uncertainty in value functions is a crucial challenge for many aspects of reinforcement learning (RL), including efficient exploration, safe decision-making, and offline RL. While deep ensembles provide a robust method for quantifying value uncertainty, they come with significant computational overhead. Single-model methods, while computationally favorable, often rely on heuristics and typically require additional propagation mechanisms for myopic uncertainty estimates. In this work we introduce universal value-function uncertainties (UVU), which, similar in spirit to random network distillation (RND), quantify uncertainty as squared prediction errors between an online learner and a fixed, randomly initialized target network. Unlike RND, UVU errors reflect policy-conditional value uncertainty, incorporating the future uncertainties any given policy may encounter. This is due to the training procedure employed in UVU: the online network is trained using temporal difference learning with a synthetic reward derived from the fixed, randomly initialized target network. We provide an extensive theoretical analysis of our approach using neural tangent kernel (NTK) theory and show that in the limit of infinite network width, UVU errors are exactly equivalent to the variance of an ensemble of independent universal value functions. Empirically, we show that UVU achieves equal performance to large ensembles on challenging multi-task offline RL settings, while offering simplicity and substantial computational savings.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21119
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Universal Value-Function Uncertainties
Zanger, Moritz A.
Weltevrede, Max
Oren, Yaniv
Van der Vaart, Pascal R.
Horsch, Caroline
Böhmer, Wendelin
Spaan, Matthijs T. J.
Machine Learning
Artificial Intelligence
Estimating epistemic uncertainty in value functions is a crucial challenge for many aspects of reinforcement learning (RL), including efficient exploration, safe decision-making, and offline RL. While deep ensembles provide a robust method for quantifying value uncertainty, they come with significant computational overhead. Single-model methods, while computationally favorable, often rely on heuristics and typically require additional propagation mechanisms for myopic uncertainty estimates. In this work we introduce universal value-function uncertainties (UVU), which, similar in spirit to random network distillation (RND), quantify uncertainty as squared prediction errors between an online learner and a fixed, randomly initialized target network. Unlike RND, UVU errors reflect policy-conditional value uncertainty, incorporating the future uncertainties any given policy may encounter. This is due to the training procedure employed in UVU: the online network is trained using temporal difference learning with a synthetic reward derived from the fixed, randomly initialized target network. We provide an extensive theoretical analysis of our approach using neural tangent kernel (NTK) theory and show that in the limit of infinite network width, UVU errors are exactly equivalent to the variance of an ensemble of independent universal value functions. Empirically, we show that UVU achieves equal performance to large ensembles on challenging multi-task offline RL settings, while offering simplicity and substantial computational savings.
title Universal Value-Function Uncertainties
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.21119