Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Zhenyu, Xie, Yuqi, Li, Jinhan, Yuan, Ye, Zhu, Yifeng, Zhu, Yuke
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914974685724672
author Jiang, Zhenyu
Xie, Yuqi
Li, Jinhan
Yuan, Ye
Zhu, Yifeng
Zhu, Yuke
author_facet Jiang, Zhenyu
Xie, Yuqi
Li, Jinhan
Yuan, Ye
Zhu, Yifeng
Zhu, Yuke
contents Humanoid robots, with their human-like embodiment, have the potential to integrate seamlessly into human environments. Critical to their coexistence and cooperation with humans is the ability to understand natural language communications and exhibit human-like behaviors. This work focuses on generating diverse whole-body motions for humanoid robots from language descriptions. We leverage human motion priors from extensive human motion datasets to initialize humanoid motions and employ the commonsense reasoning capabilities of Vision Language Models (VLMs) to edit and refine these motions. Our approach demonstrates the capability to produce natural, expressive, and text-aligned humanoid motions, validated through both simulated and real-world experiments. More videos can be found at https://ut-austin-rpl.github.io/Harmon/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12773
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions
Jiang, Zhenyu
Xie, Yuqi
Li, Jinhan
Yuan, Ye
Zhu, Yifeng
Zhu, Yuke
Robotics
Artificial Intelligence
Humanoid robots, with their human-like embodiment, have the potential to integrate seamlessly into human environments. Critical to their coexistence and cooperation with humans is the ability to understand natural language communications and exhibit human-like behaviors. This work focuses on generating diverse whole-body motions for humanoid robots from language descriptions. We leverage human motion priors from extensive human motion datasets to initialize humanoid motions and employ the commonsense reasoning capabilities of Vision Language Models (VLMs) to edit and refine these motions. Our approach demonstrates the capability to produce natural, expressive, and text-aligned humanoid motions, validated through both simulated and real-world experiments. More videos can be found at https://ut-austin-rpl.github.io/Harmon/.
title Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2410.12773