Semantic-Geometric Task Representations for Bimanual Manipulation from Human Demonstrations to Robot Action Planning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Herbert, Franziska, Prasad, Vignesh, Liu, Han, Koert, Dorothea, Chalvatzaki, Georgia
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913175411097600
author Herbert, Franziska
Prasad, Vignesh
Liu, Han
Koert, Dorothea
Chalvatzaki, Georgia
author_facet Herbert, Franziska
Prasad, Vignesh
Liu, Han
Koert, Dorothea
Chalvatzaki, Georgia
contents Learning structured task representations from human demonstrations is essential for bimanual manipulation, where action ordering, object involvement, and interaction geometry vary significantly across executions. A key challenge lies in jointly capturing the discrete semantic task structure and the temporal evolution of object-centric geometric relations in a form that supports reasoning over task progression. We introduce a semantic--geometric graph-based task representation that jointly encodes object identities, inter-object semantic relations, and per-object motion histories, via a Message Passing Neural Network (MPNN) encoder and a Transformer-based decoder. The encoder operates solely on the temporal scene graph, producing structured representations decoupled from action labels. The decoder then conditions on action-context to forecast future actions, associated objects, and object motions. This decoupling learns task-agnostic representations, enabling encoder reuse across embodiments through decoder-only finetuning on a small robot dataset. Across eleven bimanual tasks from two datasets, we find that the benefit of structured semantic--geometric representations over simpler sequence-based models grows with task variability in action ordering and object involvement. At deployment, a planner couples the action and motion predictions with learned Probabilistic Movement Primitives, achieving full task success on two real-robot bimanual tasks and outperforming graph ablations, Transformer, decoder-only, and finetuned vision-language model baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11460
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Semantic-Geometric Task Representations for Bimanual Manipulation from Human Demonstrations to Robot Action Planning
Herbert, Franziska
Prasad, Vignesh
Liu, Han
Koert, Dorothea
Chalvatzaki, Georgia
Robotics
Machine Learning
Learning structured task representations from human demonstrations is essential for bimanual manipulation, where action ordering, object involvement, and interaction geometry vary significantly across executions. A key challenge lies in jointly capturing the discrete semantic task structure and the temporal evolution of object-centric geometric relations in a form that supports reasoning over task progression. We introduce a semantic--geometric graph-based task representation that jointly encodes object identities, inter-object semantic relations, and per-object motion histories, via a Message Passing Neural Network (MPNN) encoder and a Transformer-based decoder. The encoder operates solely on the temporal scene graph, producing structured representations decoupled from action labels. The decoder then conditions on action-context to forecast future actions, associated objects, and object motions. This decoupling learns task-agnostic representations, enabling encoder reuse across embodiments through decoder-only finetuning on a small robot dataset. Across eleven bimanual tasks from two datasets, we find that the benefit of structured semantic--geometric representations over simpler sequence-based models grows with task variability in action ordering and object involvement. At deployment, a planner couples the action and motion predictions with learned Probabilistic Movement Primitives, achieving full task success on two real-robot bimanual tasks and outperforming graph ablations, Transformer, decoder-only, and finetuned vision-language model baselines.
title Semantic-Geometric Task Representations for Bimanual Manipulation from Human Demonstrations to Robot Action Planning
topic Robotics
Machine Learning
url https://arxiv.org/abs/2601.11460