RL-Driven Data Generation for Robust Vision-Based Dexterous Grasping

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kanehira, Atsushi, Wake, Naoki, Sasabuchi, Kazuhiro, Takamatsu, Jun, Ikeuchi, Katsushi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908337345396736
author Kanehira, Atsushi
Wake, Naoki
Sasabuchi, Kazuhiro
Takamatsu, Jun
Ikeuchi, Katsushi
author_facet Kanehira, Atsushi
Wake, Naoki
Sasabuchi, Kazuhiro
Takamatsu, Jun
Ikeuchi, Katsushi
contents This work presents reinforcement learning (RL)-driven data augmentation to improve the generalization of vision-action (VA) models for dexterous grasping. While real-to-sim-to-real frameworks, where a few real demonstrations seed large-scale simulated data, have proven effective for VA models, applying them to dexterous settings remains challenging: obtaining stable multi-finger contacts is nontrivial across diverse object shapes. To address this, we leverage RL to generate contact-rich grasping data across varied geometries. In line with the real-to-sim-to-real paradigm, the grasp skill is formulated as a parameterized and tunable reference trajectory refined by a residual policy learned via RL. This modular design enables trajectory-level control that is both consistent with real demonstrations and adaptable to diverse object geometries. A vision-conditioned policy trained on simulation-augmented data demonstrates strong generalization to unseen objects, highlighting the potential of our approach to alleviate the data bottleneck in training VA models.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18084
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RL-Driven Data Generation for Robust Vision-Based Dexterous Grasping
Kanehira, Atsushi
Wake, Naoki
Sasabuchi, Kazuhiro
Takamatsu, Jun
Ikeuchi, Katsushi
Robotics
This work presents reinforcement learning (RL)-driven data augmentation to improve the generalization of vision-action (VA) models for dexterous grasping. While real-to-sim-to-real frameworks, where a few real demonstrations seed large-scale simulated data, have proven effective for VA models, applying them to dexterous settings remains challenging: obtaining stable multi-finger contacts is nontrivial across diverse object shapes. To address this, we leverage RL to generate contact-rich grasping data across varied geometries. In line with the real-to-sim-to-real paradigm, the grasp skill is formulated as a parameterized and tunable reference trajectory refined by a residual policy learned via RL. This modular design enables trajectory-level control that is both consistent with real demonstrations and adaptable to diverse object geometries. A vision-conditioned policy trained on simulation-augmented data demonstrates strong generalization to unseen objects, highlighting the potential of our approach to alleviate the data bottleneck in training VA models.
title RL-Driven Data Generation for Robust Vision-Based Dexterous Grasping
topic Robotics
url https://arxiv.org/abs/2504.18084