SENSOR: Imitate Third-Person Expert's Behaviors via Active Sensoring

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Huang, Kaichen, Shao, Minghao, Wan, Shenghua, Sun, Hai-Hang, Feng, Shuai, Gan, Le, Zhan, De-Chuan
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913299339149312
author Huang, Kaichen
Shao, Minghao
Wan, Shenghua
Sun, Hai-Hang
Feng, Shuai
Gan, Le
Zhan, De-Chuan
author_facet Huang, Kaichen
Shao, Minghao
Wan, Shenghua
Sun, Hai-Hang
Feng, Shuai
Gan, Le
Zhan, De-Chuan
contents In many real-world visual Imitation Learning (IL) scenarios, there is a misalignment between the agent's and the expert's perspectives, which might lead to the failure of imitation. Previous methods have generally solved this problem by domain alignment, which incurs extra computation and storage costs, and these methods fail to handle the \textit{hard cases} where the viewpoint gap is too large. To alleviate the above problems, we introduce active sensoring in the visual IL setting and propose a model-based SENSory imitatOR (SENSOR) to automatically change the agent's perspective to match the expert's. SENSOR jointly learns a world model to capture the dynamics of latent states, a sensor policy to control the camera, and a motor policy to control the agent. Experiments on visual locomotion tasks show that SENSOR can efficiently simulate the expert's perspective and strategy, and outperforms most baseline methods.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03386
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SENSOR: Imitate Third-Person Expert's Behaviors via Active Sensoring
Huang, Kaichen
Shao, Minghao
Wan, Shenghua
Sun, Hai-Hang
Feng, Shuai
Gan, Le
Zhan, De-Chuan
Robotics
Artificial Intelligence
Machine Learning
In many real-world visual Imitation Learning (IL) scenarios, there is a misalignment between the agent's and the expert's perspectives, which might lead to the failure of imitation. Previous methods have generally solved this problem by domain alignment, which incurs extra computation and storage costs, and these methods fail to handle the \textit{hard cases} where the viewpoint gap is too large. To alleviate the above problems, we introduce active sensoring in the visual IL setting and propose a model-based SENSory imitatOR (SENSOR) to automatically change the agent's perspective to match the expert's. SENSOR jointly learns a world model to capture the dynamics of latent states, a sensor policy to control the camera, and a motor policy to control the agent. Experiments on visual locomotion tasks show that SENSOR can efficiently simulate the expert's perspective and strategy, and outperforms most baseline methods.
title SENSOR: Imitate Third-Person Expert's Behaviors via Active Sensoring
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.03386