Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: An, Boshi, Yang, Chenyu, Katzschmann, Robert
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914122086481920
author An, Boshi
Yang, Chenyu
Katzschmann, Robert
author_facet An, Boshi
Yang, Chenyu
Katzschmann, Robert
contents We adapt a pre-trained Vision-Language-Action (VLA) model (Open-VLA) for dexterous human-robot collaboration with minimal language prompting. Our approach adds (i) FiLM conditioning to visual backbones for task-aware perception, (ii) an auxiliary intent head that predicts collaborator hand pose and target cues, and (iii) action-space post-processing that predicts compact deltas (position/rotation) and PCA-reduced finger joints before mapping to full commands. Using a multi-view, teleoperated Franka and Mimic-hand dataset augmented with MediaPipe hand poses, we demonstrate that delta actions are well-behaved and that four principal components explain ~96% of hand-joint variance. Ablations identify action post-processing as the primary performance driver; auxiliary intent helps, FiLM is mixed, and a directional motion loss is detrimental. A real-time stack (~0.3 s latency on one RTX 4090) composes "pick-up" and "pass" into a long-horizon behavior. We surface "trainer overfitting" to specific demonstrators as the key limitation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models
An, Boshi
Yang, Chenyu
Katzschmann, Robert
Robotics
We adapt a pre-trained Vision-Language-Action (VLA) model (Open-VLA) for dexterous human-robot collaboration with minimal language prompting. Our approach adds (i) FiLM conditioning to visual backbones for task-aware perception, (ii) an auxiliary intent head that predicts collaborator hand pose and target cues, and (iii) action-space post-processing that predicts compact deltas (position/rotation) and PCA-reduced finger joints before mapping to full commands. Using a multi-view, teleoperated Franka and Mimic-hand dataset augmented with MediaPipe hand poses, we demonstrate that delta actions are well-behaved and that four principal components explain ~96% of hand-joint variance. Ablations identify action post-processing as the primary performance driver; auxiliary intent helps, FiLM is mixed, and a directional motion loss is detrimental. A real-time stack (~0.3 s latency on one RTX 4090) composes "pick-up" and "pass" into a long-horizon behavior. We surface "trainer overfitting" to specific demonstrators as the key limitation.
title Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models
topic Robotics
url https://arxiv.org/abs/2510.25713