Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dupuis, Tom, Rabarisoa, Jaonary, Pham, Quoc-Cuong, Filliat, David
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:https://arxiv.org/abs/2306.08537
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913199237890048
author Dupuis, Tom
Rabarisoa, Jaonary
Pham, Quoc-Cuong
Filliat, David
author_facet Dupuis, Tom
Rabarisoa, Jaonary
Pham, Quoc-Cuong
Filliat, David
contents End-to-end reinforcement learning on images showed significant progress in the recent years. Data-based approach leverage data augmentation and domain randomization while representation learning methods use auxiliary losses to learn task-relevant features. Yet, reinforcement still struggles in visually diverse environments full of distractions and spurious noise. In this work, we tackle the problem of robust visual control at its core and present VIBR (View-Invariant Bellman Residuals), a method that combines multi-view training and invariant prediction to reduce out-of-distribution (OOD) generalization gap for RL based visuomotor control. Our model-free approach improve baselines performances without the need of additional representation learning objectives and with limited additional computational cost. We show that VIBR outperforms existing methods on complex visuo-motor control environment with high visual perturbation. Our approach achieves state-of the-art results on the Distracting Control Suite benchmark, a challenging benchmark still not solved by current methods, where we evaluate the robustness to a number of visual perturbators, as well as OOD generalization and extrapolation capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2306_08537
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VIBR: Learning View-Invariant Value Functions for Robust Visual Control
Dupuis, Tom
Rabarisoa, Jaonary
Pham, Quoc-Cuong
Filliat, David
Machine Learning
Computer Vision and Pattern Recognition
Systems and Control
End-to-end reinforcement learning on images showed significant progress in the recent years. Data-based approach leverage data augmentation and domain randomization while representation learning methods use auxiliary losses to learn task-relevant features. Yet, reinforcement still struggles in visually diverse environments full of distractions and spurious noise. In this work, we tackle the problem of robust visual control at its core and present VIBR (View-Invariant Bellman Residuals), a method that combines multi-view training and invariant prediction to reduce out-of-distribution (OOD) generalization gap for RL based visuomotor control. Our model-free approach improve baselines performances without the need of additional representation learning objectives and with limited additional computational cost. We show that VIBR outperforms existing methods on complex visuo-motor control environment with high visual perturbation. Our approach achieves state-of the-art results on the Distracting Control Suite benchmark, a challenging benchmark still not solved by current methods, where we evaluate the robustness to a number of visual perturbators, as well as OOD generalization and extrapolation capabilities.
title VIBR: Learning View-Invariant Value Functions for Robust Visual Control
topic Machine Learning
Computer Vision and Pattern Recognition
Systems and Control
url https://arxiv.org/abs/2306.08537