Maximum diffusion reinforcement learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Berrueta, Thomas A., Pinosky, Allison, Murphey, Todd D.
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914810758692864
author Berrueta, Thomas A.
Pinosky, Allison
Murphey, Todd D.
author_facet Berrueta, Thomas A.
Pinosky, Allison
Murphey, Todd D.
contents Robots and animals both experience the world through their bodies and senses. Their embodiment constrains their experiences, ensuring they unfold continuously in space and time. As a result, the experiences of embodied agents are intrinsically correlated. Correlations create fundamental challenges for machine learning, as most techniques rely on the assumption that data are independent and identically distributed. In reinforcement learning, where data are directly collected from an agent's sequential experiences, violations of this assumption are often unavoidable. Here, we derive a method that overcomes this issue by exploiting the statistical mechanics of ergodic processes, which we term maximum diffusion reinforcement learning. By decorrelating agent experiences, our approach provably enables single-shot learning in continuous deployments over the course of individual task attempts. Moreover, we prove our approach generalizes well-known maximum entropy techniques, and robustly exceeds state-of-the-art performance across popular benchmarks. Our results at the nexus of physics, learning, and control form a foundation for transparent and reliable decision-making in embodied reinforcement learning agents.
format Preprint
id arxiv_https___arxiv_org_abs_2309_15293
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Maximum diffusion reinforcement learning
Berrueta, Thomas A.
Pinosky, Allison
Murphey, Todd D.
Machine Learning
Statistical Mechanics
Artificial Intelligence
Robotics
Robots and animals both experience the world through their bodies and senses. Their embodiment constrains their experiences, ensuring they unfold continuously in space and time. As a result, the experiences of embodied agents are intrinsically correlated. Correlations create fundamental challenges for machine learning, as most techniques rely on the assumption that data are independent and identically distributed. In reinforcement learning, where data are directly collected from an agent's sequential experiences, violations of this assumption are often unavoidable. Here, we derive a method that overcomes this issue by exploiting the statistical mechanics of ergodic processes, which we term maximum diffusion reinforcement learning. By decorrelating agent experiences, our approach provably enables single-shot learning in continuous deployments over the course of individual task attempts. Moreover, we prove our approach generalizes well-known maximum entropy techniques, and robustly exceeds state-of-the-art performance across popular benchmarks. Our results at the nexus of physics, learning, and control form a foundation for transparent and reliable decision-making in embodied reinforcement learning agents.
title Maximum diffusion reinforcement learning
topic Machine Learning
Statistical Mechanics
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2309.15293