DeGuV: Depth-Guided Visual Reinforcement Learning for Generalization and Interpretability in Manipulation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Pham, Tien, Chi, Xinyun, Nguyen, Khang, Huber, Manfred, Cangelosi, Angelo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911139417292800
author Pham, Tien
Chi, Xinyun
Nguyen, Khang
Huber, Manfred
Cangelosi, Angelo
author_facet Pham, Tien
Chi, Xinyun
Nguyen, Khang
Huber, Manfred
Cangelosi, Angelo
contents Reinforcement learning (RL) agents can learn to solve complex tasks from visual inputs, but generalizing these learned skills to new environments remains a major challenge in RL application, especially robotics. While data augmentation can improve generalization, it often compromises sample efficiency and training stability. This paper introduces DeGuV, an RL framework that enhances both generalization and sample efficiency. In specific, we leverage a learnable masker network that produces a mask from the depth input, preserving only critical visual information while discarding irrelevant pixels. Through this, we ensure that our RL agents focus on essential features, improving robustness under data augmentation. In addition, we incorporate contrastive learning and stabilize Q-value estimation under augmentation to further enhance sample efficiency and training stability. We evaluate our proposed method on the RL-ViGen benchmark using the Franka Emika robot and demonstrate its effectiveness in zero-shot sim-to-real transfer. Our results show that DeGuV outperforms state-of-the-art methods in both generalization and sample efficiency while also improving interpretability by highlighting the most relevant regions in the visual input
format Preprint
id arxiv_https___arxiv_org_abs_2509_04970
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeGuV: Depth-Guided Visual Reinforcement Learning for Generalization and Interpretability in Manipulation
Pham, Tien
Chi, Xinyun
Nguyen, Khang
Huber, Manfred
Cangelosi, Angelo
Robotics
Artificial Intelligence
Reinforcement learning (RL) agents can learn to solve complex tasks from visual inputs, but generalizing these learned skills to new environments remains a major challenge in RL application, especially robotics. While data augmentation can improve generalization, it often compromises sample efficiency and training stability. This paper introduces DeGuV, an RL framework that enhances both generalization and sample efficiency. In specific, we leverage a learnable masker network that produces a mask from the depth input, preserving only critical visual information while discarding irrelevant pixels. Through this, we ensure that our RL agents focus on essential features, improving robustness under data augmentation. In addition, we incorporate contrastive learning and stabilize Q-value estimation under augmentation to further enhance sample efficiency and training stability. We evaluate our proposed method on the RL-ViGen benchmark using the Franka Emika robot and demonstrate its effectiveness in zero-shot sim-to-real transfer. Our results show that DeGuV outperforms state-of-the-art methods in both generalization and sample efficiency while also improving interpretability by highlighting the most relevant regions in the visual input
title DeGuV: Depth-Guided Visual Reinforcement Learning for Generalization and Interpretability in Manipulation
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2509.04970