Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Zekai, Shi, Ye, Ji, Kaiyang, Xu, Lan, Huang, Shaoli, Wang, Jingya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914368307855360
author Deng, Zekai
Shi, Ye
Ji, Kaiyang
Xu, Lan
Huang, Shaoli
Wang, Jingya
author_facet Deng, Zekai
Shi, Ye
Ji, Kaiyang
Xu, Lan
Huang, Shaoli
Wang, Jingya
contents Human-object interaction (HOI) synthesis is crucial for applications in animation, simulation, and robotics. However, existing approaches either rely on expensive motion capture data or require manual reward engineering, limiting their scalability and generalizability. In this work, we introduce the first unified physics-based HOI framework that leverages Vision-Language Models (VLMs) to enable long-horizon interactions with diverse object types, including static, dynamic, and articulated objects. We introduce VLM-Guided Relative Movement Dynamics (RMD), a fine-grained spatio-temporal bipartite representation that automatically constructs goal states and reward functions for reinforcement learning. By encoding structured relationships between human and object parts, RMD enables VLMs to generate semantically grounded, interaction-aware motion guidance without manual reward tuning. To support our methodology, we present Interplay, a novel dataset with thousands of long-horizon static and dynamic interaction plans. Extensive experiments demonstrate that our framework outperforms existing methods in synthesizing natural, human-like motions across both simple single-task and complex multi-task scenarios. For more details, please refer to our project webpage: https://vlm-rmd.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18349
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy
Deng, Zekai
Shi, Ye
Ji, Kaiyang
Xu, Lan
Huang, Shaoli
Wang, Jingya
Computer Vision and Pattern Recognition
Human-object interaction (HOI) synthesis is crucial for applications in animation, simulation, and robotics. However, existing approaches either rely on expensive motion capture data or require manual reward engineering, limiting their scalability and generalizability. In this work, we introduce the first unified physics-based HOI framework that leverages Vision-Language Models (VLMs) to enable long-horizon interactions with diverse object types, including static, dynamic, and articulated objects. We introduce VLM-Guided Relative Movement Dynamics (RMD), a fine-grained spatio-temporal bipartite representation that automatically constructs goal states and reward functions for reinforcement learning. By encoding structured relationships between human and object parts, RMD enables VLMs to generate semantically grounded, interaction-aware motion guidance without manual reward tuning. To support our methodology, we present Interplay, a novel dataset with thousands of long-horizon static and dynamic interaction plans. Extensive experiments demonstrate that our framework outperforms existing methods in synthesizing natural, human-like motions across both simple single-task and complex multi-task scenarios. For more details, please refer to our project webpage: https://vlm-rmd.github.io/.
title Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18349