PoseLess: Depth-Free Vision-to-Joint Control via Direct Image Mapping with VLM

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dao, Alan, Vu, Dinh Bach, Anh, Tuan Le Duc, Huy, Bui Quang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916649423077376
author Dao, Alan
Vu, Dinh Bach
Anh, Tuan Le Duc
Huy, Bui Quang
author_facet Dao, Alan
Vu, Dinh Bach
Anh, Tuan Le Duc
Huy, Bui Quang
contents This paper introduces PoseLess, a novel framework for robot hand control that eliminates the need for explicit pose estimation by directly mapping 2D images to joint angles using projected representations. Our approach leverages synthetic training data generated through randomized joint configurations, enabling zero-shot generalization to real-world scenarios and cross-morphology transfer from robotic to human hands. By projecting visual inputs and employing a transformer-based decoder, PoseLess achieves robust, low-latency control while addressing challenges such as depth ambiguity and data scarcity. Experimental results demonstrate competitive performance in joint angle prediction accuracy without relying on any human-labelled dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2503_07111
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PoseLess: Depth-Free Vision-to-Joint Control via Direct Image Mapping with VLM
Dao, Alan
Vu, Dinh Bach
Anh, Tuan Le Duc
Huy, Bui Quang
Robotics
Computation and Language
This paper introduces PoseLess, a novel framework for robot hand control that eliminates the need for explicit pose estimation by directly mapping 2D images to joint angles using projected representations. Our approach leverages synthetic training data generated through randomized joint configurations, enabling zero-shot generalization to real-world scenarios and cross-morphology transfer from robotic to human hands. By projecting visual inputs and employing a transformer-based decoder, PoseLess achieves robust, low-latency control while addressing challenges such as depth ambiguity and data scarcity. Experimental results demonstrate competitive performance in joint angle prediction accuracy without relying on any human-labelled dataset.
title PoseLess: Depth-Free Vision-to-Joint Control via Direct Image Mapping with VLM
topic Robotics
Computation and Language
url https://arxiv.org/abs/2503.07111