Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918132346519552 |
|---|---|
| author | Almuzairee, Abdulaziz Patil, Rohan Bhatt, Dwait Christensen, Henrik I. |
| author_facet | Almuzairee, Abdulaziz Patil, Rohan Bhatt, Dwait Christensen, Henrik I. |
| contents | Vision is well-known for its use in manipulation, especially using visual servoing. Due to the 3D nature of the world, using multiple camera views and merging them creates better representations for Q-learning and in turn, trains more sample efficient policies. Nevertheless, these multi-view policies are sensitive to failing cameras and can be burdensome to deploy. To mitigate these issues, we introduce a Merge And Disentanglement (MAD) algorithm that efficiently merges views to increase sample efficiency while simultaneously disentangling views by augmenting multi-view feature inputs with single-view features. This produces robust policies and allows lightweight deployment. We demonstrate the efficiency and robustness of our approach using Meta-World and ManiSkill3. For project website and code, see https://aalmuzairee.github.io/mad |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_04619 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation Almuzairee, Abdulaziz Patil, Rohan Bhatt, Dwait Christensen, Henrik I. Machine Learning Computer Vision and Pattern Recognition Robotics Vision is well-known for its use in manipulation, especially using visual servoing. Due to the 3D nature of the world, using multiple camera views and merging them creates better representations for Q-learning and in turn, trains more sample efficient policies. Nevertheless, these multi-view policies are sensitive to failing cameras and can be burdensome to deploy. To mitigate these issues, we introduce a Merge And Disentanglement (MAD) algorithm that efficiently merges views to increase sample efficiency while simultaneously disentangling views by augmenting multi-view feature inputs with single-view features. This produces robust policies and allows lightweight deployment. We demonstrate the efficiency and robustness of our approach using Meta-World and ManiSkill3. For project website and code, see https://aalmuzairee.github.io/mad |
| title | Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation |
| topic | Machine Learning Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2505.04619 |