Coordinated Humanoid Manipulation with Choice Policies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qi, Haozhi, Wang, Yen-Jen, Lin, Toru, Yi, Brent, Ma, Yi, Sreenath, Koushil, Malik, Jitendra
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917178606878720
author Qi, Haozhi
Wang, Yen-Jen
Lin, Toru
Yi, Brent
Ma, Yi
Sreenath, Koushil
Malik, Jitendra
author_facet Qi, Haozhi
Wang, Yen-Jen
Lin, Toru
Yi, Brent
Ma, Yi
Sreenath, Koushil
Malik, Jitendra
contents Humanoid robots hold great promise for operating in human-centric environments, yet achieving robust whole-body coordination across the head, hands, and legs remains a major challenge. We present a system that combines a modular teleoperation interface with a scalable learning framework to address this problem. Our teleoperation design decomposes humanoid control into intuitive submodules, which include hand-eye coordination, grasp primitives, arm end-effector tracking, and locomotion. This modularity allows us to collect high-quality demonstrations efficiently. Building on this, we introduce Choice Policy, an imitation learning approach that generates multiple candidate actions and learns to score them. This architecture enables both fast inference and effective modeling of multimodal behaviors. We validate our approach on two real-world tasks: dishwasher loading and whole-body loco-manipulation for whiteboard wiping. Experiments show that Choice Policy significantly outperforms diffusion policies and standard behavior cloning. Furthermore, our results indicate that hand-eye coordination is critical for success in long-horizon tasks. Our work demonstrates a practical path toward scalable data collection and learning for coordinated humanoid manipulation in unstructured environments.
format Preprint
id arxiv_https___arxiv_org_abs_2512_25072
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Coordinated Humanoid Manipulation with Choice Policies
Qi, Haozhi
Wang, Yen-Jen
Lin, Toru
Yi, Brent
Ma, Yi
Sreenath, Koushil
Malik, Jitendra
Robotics
Artificial Intelligence
Machine Learning
Humanoid robots hold great promise for operating in human-centric environments, yet achieving robust whole-body coordination across the head, hands, and legs remains a major challenge. We present a system that combines a modular teleoperation interface with a scalable learning framework to address this problem. Our teleoperation design decomposes humanoid control into intuitive submodules, which include hand-eye coordination, grasp primitives, arm end-effector tracking, and locomotion. This modularity allows us to collect high-quality demonstrations efficiently. Building on this, we introduce Choice Policy, an imitation learning approach that generates multiple candidate actions and learns to score them. This architecture enables both fast inference and effective modeling of multimodal behaviors. We validate our approach on two real-world tasks: dishwasher loading and whole-body loco-manipulation for whiteboard wiping. Experiments show that Choice Policy significantly outperforms diffusion policies and standard behavior cloning. Furthermore, our results indicate that hand-eye coordination is critical for success in long-horizon tasks. Our work demonstrates a practical path toward scalable data collection and learning for coordinated humanoid manipulation in unstructured environments.
title Coordinated Humanoid Manipulation with Choice Policies
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.25072