Saved in:
Bibliographic Details
Main Authors: Nedergaard, Alexander, Morales, Pablo A.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.06355
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908971750653952
author Nedergaard, Alexander
Morales, Pablo A.
author_facet Nedergaard, Alexander
Morales, Pablo A.
contents Learning in environments with sparse rewards remains a fundamental challenge in reinforcement learning. Artificial curiosity addresses this limitation through intrinsic rewards to guide exploration, however, the precise formulation of these rewards has remained elusive. Ideally, such rewards should depend on the agent's information about the environment, remaining agnostic to its representation -- an invariance central to information geometry. Leveraging this, we show that information monotonicity and invariance under the agent-environment interaction uniquely constrains intrinsic rewards to strictly concave functions of the reciprocal occupancy. Requiring these rewards to yield a principled exploration-exploitation trade-off, via information geodesic interpolation on the occupancy manifold, effectively limits the candidates to those determined by a scalar parameter. Remarkably, special values of this parameter are found to correspond to count-based and maximum entropy exploration. This framework provides important constraints to the engineering of intrinsic rewards while integrating foundational exploration methods into a single, cohesive model.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06355
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Information-Geometric Approach to Artificial Curiosity
Nedergaard, Alexander
Morales, Pablo A.
Machine Learning
Learning in environments with sparse rewards remains a fundamental challenge in reinforcement learning. Artificial curiosity addresses this limitation through intrinsic rewards to guide exploration, however, the precise formulation of these rewards has remained elusive. Ideally, such rewards should depend on the agent's information about the environment, remaining agnostic to its representation -- an invariance central to information geometry. Leveraging this, we show that information monotonicity and invariance under the agent-environment interaction uniquely constrains intrinsic rewards to strictly concave functions of the reciprocal occupancy. Requiring these rewards to yield a principled exploration-exploitation trade-off, via information geodesic interpolation on the occupancy manifold, effectively limits the candidates to those determined by a scalar parameter. Remarkably, special values of this parameter are found to correspond to count-based and maximum entropy exploration. This framework provides important constraints to the engineering of intrinsic rewards while integrating foundational exploration methods into a single, cohesive model.
title An Information-Geometric Approach to Artificial Curiosity
topic Machine Learning
url https://arxiv.org/abs/2504.06355