Generative Simulation for Policy Learning in Physical Human-Robot Interaction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Junxiang, Xu, Xinwen, Wu, Tiancheng, Millan, Julian, Pechuk, Nir, Erickson, Zackory
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908950598778880
author Wang, Junxiang
Xu, Xinwen
Wu, Tiancheng
Millan, Julian
Pechuk, Nir
Erickson, Zackory
author_facet Wang, Junxiang
Xu, Xinwen
Wu, Tiancheng
Millan, Julian
Pechuk, Nir
Erickson, Zackory
contents Developing autonomous physical human-robot interaction (pHRI) systems is limited by the scarcity of large-scale training data to learn robust robot behaviors for real-world applications. In this paper, we introduce a zero-shot "text2sim2real" generative simulation framework that automatically synthesizes diverse pHRI scenarios from high-level natural-language prompts. Leveraging Large Language Models (LLMs) and Vision-Language Models (VLMs), our pipeline procedurally generates soft-body human models, scene layouts, and robot motion trajectories for assistive tasks. We utilize this framework to autonomously collect large-scale synthetic demonstration datasets and then train vision-based imitation learning policies operating on segmented point clouds. We evaluate our approach through a user study on two physically assistive tasks: scratching and bathing. Our learned policies successfully achieve zero-shot sim-to-real transfer, attaining success rates exceeding 80% and demonstrating resilience to unscripted human motion. Overall, we introduce the first generative simulation pipeline for pHRI applications, automating simulation environment synthesis, data collection, and policy learning. Additional information may be found on our project website: https://rchi-lab.github.io/gen_phri/
format Preprint
id arxiv_https___arxiv_org_abs_2604_08664
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Generative Simulation for Policy Learning in Physical Human-Robot Interaction
Wang, Junxiang
Xu, Xinwen
Wu, Tiancheng
Millan, Julian
Pechuk, Nir
Erickson, Zackory
Robotics
Developing autonomous physical human-robot interaction (pHRI) systems is limited by the scarcity of large-scale training data to learn robust robot behaviors for real-world applications. In this paper, we introduce a zero-shot "text2sim2real" generative simulation framework that automatically synthesizes diverse pHRI scenarios from high-level natural-language prompts. Leveraging Large Language Models (LLMs) and Vision-Language Models (VLMs), our pipeline procedurally generates soft-body human models, scene layouts, and robot motion trajectories for assistive tasks. We utilize this framework to autonomously collect large-scale synthetic demonstration datasets and then train vision-based imitation learning policies operating on segmented point clouds. We evaluate our approach through a user study on two physically assistive tasks: scratching and bathing. Our learned policies successfully achieve zero-shot sim-to-real transfer, attaining success rates exceeding 80% and demonstrating resilience to unscripted human motion. Overall, we introduce the first generative simulation pipeline for pHRI applications, automating simulation environment synthesis, data collection, and policy learning. Additional information may be found on our project website: https://rchi-lab.github.io/gen_phri/
title Generative Simulation for Policy Learning in Physical Human-Robot Interaction
topic Robotics
url https://arxiv.org/abs/2604.08664