Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Almuzairee, Abdulaziz, Patil, Rohan, Bhatt, Dwait, Christensen, Henrik I.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918132346519552
author Almuzairee, Abdulaziz
Patil, Rohan
Bhatt, Dwait
Christensen, Henrik I.
author_facet Almuzairee, Abdulaziz
Patil, Rohan
Bhatt, Dwait
Christensen, Henrik I.
contents Vision is well-known for its use in manipulation, especially using visual servoing. Due to the 3D nature of the world, using multiple camera views and merging them creates better representations for Q-learning and in turn, trains more sample efficient policies. Nevertheless, these multi-view policies are sensitive to failing cameras and can be burdensome to deploy. To mitigate these issues, we introduce a Merge And Disentanglement (MAD) algorithm that efficiently merges views to increase sample efficiency while simultaneously disentangling views by augmenting multi-view feature inputs with single-view features. This produces robust policies and allows lightweight deployment. We demonstrate the efficiency and robustness of our approach using Meta-World and ManiSkill3. For project website and code, see https://aalmuzairee.github.io/mad
format Preprint
id arxiv_https___arxiv_org_abs_2505_04619
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation
Almuzairee, Abdulaziz
Patil, Rohan
Bhatt, Dwait
Christensen, Henrik I.
Machine Learning
Computer Vision and Pattern Recognition
Robotics
Vision is well-known for its use in manipulation, especially using visual servoing. Due to the 3D nature of the world, using multiple camera views and merging them creates better representations for Q-learning and in turn, trains more sample efficient policies. Nevertheless, these multi-view policies are sensitive to failing cameras and can be burdensome to deploy. To mitigate these issues, we introduce a Merge And Disentanglement (MAD) algorithm that efficiently merges views to increase sample efficiency while simultaneously disentangling views by augmenting multi-view feature inputs with single-view features. This produces robust policies and allows lightweight deployment. We demonstrate the efficiency and robustness of our approach using Meta-World and ManiSkill3. For project website and code, see https://aalmuzairee.github.io/mad
title Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation
topic Machine Learning
Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2505.04619