Bridging the Sim-to-Real Gap from the Information Bottleneck Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Haoran, Wu, Peilin, Bai, Chenjia, Lai, Hang, Wang, Lingxiao, Pan, Ling, Hu, Xiaolin, Zhang, Weinan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909346986721280
author He, Haoran
Wu, Peilin
Bai, Chenjia
Lai, Hang
Wang, Lingxiao
Pan, Ling
Hu, Xiaolin
Zhang, Weinan
author_facet He, Haoran
Wu, Peilin
Bai, Chenjia
Lai, Hang
Wang, Lingxiao
Pan, Ling
Hu, Xiaolin
Zhang, Weinan
contents Reinforcement Learning (RL) has recently achieved remarkable success in robotic control. However, most works in RL operate in simulated environments where privileged knowledge (e.g., dynamics, surroundings, terrains) is readily available. Conversely, in real-world scenarios, robot agents usually rely solely on local states (e.g., proprioceptive feedback of robot joints) to select actions, leading to a significant sim-to-real gap. Existing methods address this gap by either gradually reducing the reliance on privileged knowledge or performing a two-stage policy imitation. However, we argue that these methods are limited in their ability to fully leverage the available privileged knowledge, resulting in suboptimal performance. In this paper, we formulate the sim-to-real gap as an information bottleneck problem and therefore propose a novel privileged knowledge distillation method called the Historical Information Bottleneck (HIB). In particular, HIB learns a privileged knowledge representation from historical trajectories by capturing the underlying changeable dynamic information. Theoretical analysis shows that the learned privileged knowledge representation helps reduce the value discrepancy between the oracle and learned policies. Empirical experiments on both simulated and real-world tasks demonstrate that HIB yields improved generalizability compared to previous methods. Videos of real-world experiments are available at https://sites.google.com/view/history-ib .
format Preprint
id arxiv_https___arxiv_org_abs_2305_18464
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Bridging the Sim-to-Real Gap from the Information Bottleneck Perspective
He, Haoran
Wu, Peilin
Bai, Chenjia
Lai, Hang
Wang, Lingxiao
Pan, Ling
Hu, Xiaolin
Zhang, Weinan
Machine Learning
Robotics
Reinforcement Learning (RL) has recently achieved remarkable success in robotic control. However, most works in RL operate in simulated environments where privileged knowledge (e.g., dynamics, surroundings, terrains) is readily available. Conversely, in real-world scenarios, robot agents usually rely solely on local states (e.g., proprioceptive feedback of robot joints) to select actions, leading to a significant sim-to-real gap. Existing methods address this gap by either gradually reducing the reliance on privileged knowledge or performing a two-stage policy imitation. However, we argue that these methods are limited in their ability to fully leverage the available privileged knowledge, resulting in suboptimal performance. In this paper, we formulate the sim-to-real gap as an information bottleneck problem and therefore propose a novel privileged knowledge distillation method called the Historical Information Bottleneck (HIB). In particular, HIB learns a privileged knowledge representation from historical trajectories by capturing the underlying changeable dynamic information. Theoretical analysis shows that the learned privileged knowledge representation helps reduce the value discrepancy between the oracle and learned policies. Empirical experiments on both simulated and real-world tasks demonstrate that HIB yields improved generalizability compared to previous methods. Videos of real-world experiments are available at https://sites.google.com/view/history-ib .
title Bridging the Sim-to-Real Gap from the Information Bottleneck Perspective
topic Machine Learning
Robotics
url https://arxiv.org/abs/2305.18464