To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Yuda, Rohatgi, Dhruv, Singh, Aarti, Bagnell, J. Andrew
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912626576982016
author Song, Yuda
Rohatgi, Dhruv
Singh, Aarti
Bagnell, J. Andrew
author_facet Song, Yuda
Rohatgi, Dhruv
Singh, Aarti
Bagnell, J. Andrew
contents Partial observability is a notorious challenge in reinforcement learning (RL), due to the need to learn complex, history-dependent policies. Recent empirical successes have used privileged expert distillation--which leverages availability of latent state information during training (e.g., from a simulator) to learn and imitate the optimal latent, Markovian policy--to disentangle the task of "learning to see" from "learning to act". While expert distillation is more computationally efficient than RL without latent state information, it also has well-documented failure modes. In this paper--through a simple but instructive theoretical model called the perturbed Block MDP, and controlled experiments on challenging simulated locomotion tasks--we investigate the algorithmic trade-off between privileged expert distillation and standard RL without privileged information. Our main findings are: (1) The trade-off empirically hinges on the stochasticity of the latent dynamics, as theoretically predicted by contrasting approximate decodability with belief contraction in the perturbed Block MDP; and (2) The optimal latent policy is not always the best latent policy to distill. Our results suggest new guidelines for effectively exploiting privileged information, potentially advancing the efficiency of policy learning across many practical partially observable domains.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03207
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning
Song, Yuda
Rohatgi, Dhruv
Singh, Aarti
Bagnell, J. Andrew
Machine Learning
Partial observability is a notorious challenge in reinforcement learning (RL), due to the need to learn complex, history-dependent policies. Recent empirical successes have used privileged expert distillation--which leverages availability of latent state information during training (e.g., from a simulator) to learn and imitate the optimal latent, Markovian policy--to disentangle the task of "learning to see" from "learning to act". While expert distillation is more computationally efficient than RL without latent state information, it also has well-documented failure modes. In this paper--through a simple but instructive theoretical model called the perturbed Block MDP, and controlled experiments on challenging simulated locomotion tasks--we investigate the algorithmic trade-off between privileged expert distillation and standard RL without privileged information. Our main findings are: (1) The trade-off empirically hinges on the stochasticity of the latent dynamics, as theoretically predicted by contrasting approximate decodability with belief contraction in the perturbed Block MDP; and (2) The optimal latent policy is not always the best latent policy to distill. Our results suggest new guidelines for effectively exploiting privileged information, potentially advancing the efficiency of policy learning across many practical partially observable domains.
title To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2510.03207