SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911703835344896 |
|---|---|
| author | Luo, Zhengyi Yuan, Ye Wang, Tingwu Li, Chenran Castañeda, Fernando Chen, Sirui Cao, Zi-Ang Li, Jiefeng Minor, David Ben, Qingwei Park, Jinhyung Sami, David Wang, Zi Da, Xingye Ding, Runyu Hogg, Cyrus Song, Lina Lim, Edy Jeong, Eugene He, Tairan Xue, Haoru Xiao, Wenli Yuen, Simon Kautz, Jan Chang, Yan Iqbal, Umar Fan, Linxi "Jim" Zhu, Yuke |
| author_facet | Luo, Zhengyi Yuan, Ye Wang, Tingwu Li, Chenran Castañeda, Fernando Chen, Sirui Cao, Zi-Ang Li, Jiefeng Minor, David Ben, Qingwei Park, Jinhyung Sami, David Wang, Zi Da, Xingye Ding, Runyu Hogg, Cyrus Song, Lina Lim, Edy Jeong, Eugene He, Tairan Xue, Haoru Xiao, Wenli Yuen, Simon Kautz, Jan Chang, Yan Iqbal, Umar Fan, Linxi "Jim" Zhu, Yuke |
| contents | Despite the rise of billion-parameter foundation models trained across thousands of GPUs, similar scaling gains have not been shown for humanoid control. Current neural controllers for humanoids remain modest in size, target a limited set of behaviors, and are trained on a handful of GPUs. We show that scaling model capacity, data, and compute yields a generalist humanoid controller capable of natural, robust whole-body movements. We position motion tracking as a scalable task for humanoid control, leveraging dense supervision from diverse motion-capture data to acquire human motion priors without manual reward engineering. We build a foundation model for motion tracking by scaling along three axes: network size (1.2M to 42M parameters), dataset volume (100M+ frames from 700 hours of motion capture), and compute (21k GPU hours). Beyond demonstrating the benefits of scale, we further show downstream utility through: (1) a real-time kinematic planner bridging motion tracking to tasks such as navigation, enabling natural and interactive control, and (2) a unified token space supporting VR teleoperation and vision-language-action (VLA) models with a single policy. Through this interface, we demonstrate autonomous VLA-driven whole-body loco-manipulation requiring coordinated hand and foot placement. Scaling motion tracking exhibits favorable properties: performance improves steadily with compute and data diversity, and learned policies generalize to unseen motions, establishing motion tracking at scale as a practical foundation for humanoid control. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_07820 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control Luo, Zhengyi Yuan, Ye Wang, Tingwu Li, Chenran Castañeda, Fernando Chen, Sirui Cao, Zi-Ang Li, Jiefeng Minor, David Ben, Qingwei Park, Jinhyung Sami, David Wang, Zi Da, Xingye Ding, Runyu Hogg, Cyrus Song, Lina Lim, Edy Jeong, Eugene He, Tairan Xue, Haoru Xiao, Wenli Yuen, Simon Kautz, Jan Chang, Yan Iqbal, Umar Fan, Linxi "Jim" Zhu, Yuke Robotics Artificial Intelligence Computer Vision and Pattern Recognition Graphics Systems and Control Despite the rise of billion-parameter foundation models trained across thousands of GPUs, similar scaling gains have not been shown for humanoid control. Current neural controllers for humanoids remain modest in size, target a limited set of behaviors, and are trained on a handful of GPUs. We show that scaling model capacity, data, and compute yields a generalist humanoid controller capable of natural, robust whole-body movements. We position motion tracking as a scalable task for humanoid control, leveraging dense supervision from diverse motion-capture data to acquire human motion priors without manual reward engineering. We build a foundation model for motion tracking by scaling along three axes: network size (1.2M to 42M parameters), dataset volume (100M+ frames from 700 hours of motion capture), and compute (21k GPU hours). Beyond demonstrating the benefits of scale, we further show downstream utility through: (1) a real-time kinematic planner bridging motion tracking to tasks such as navigation, enabling natural and interactive control, and (2) a unified token space supporting VR teleoperation and vision-language-action (VLA) models with a single policy. Through this interface, we demonstrate autonomous VLA-driven whole-body loco-manipulation requiring coordinated hand and foot placement. Scaling motion tracking exhibits favorable properties: performance improves steadily with compute and data diversity, and learned policies generalize to unseen motions, establishing motion tracking at scale as a practical foundation for humanoid control. |
| title | SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control |
| topic | Robotics Artificial Intelligence Computer Vision and Pattern Recognition Graphics Systems and Control |
| url | https://arxiv.org/abs/2511.07820 |