Hierarchical World Models as Visual Whole-Body Humanoid Controllers
Fuente:
arXiv
Saved in:
| Main Authors: | Hansen, Nicklas, S V, Jyothir, Sobal, Vlad, LeCun, Yann, Wang, Xiaolong, Su, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Massively Multitask World Models for Continuous Control
by: Hansen, Nicklas, et al.
Published: (2025)
by: Hansen, Nicklas, et al.
Published: (2025)
TD-MPC2: Scalable, Robust World Models for Continuous Control
by: Hansen, Nicklas, et al.
Published: (2023)
by: Hansen, Nicklas, et al.
Published: (2023)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation
by: He, Xialin, et al.
Published: (2026)
by: He, Xialin, et al.
Published: (2026)
Visual Whole-Body Control for Legged Loco-Manipulation
by: Liu, Minghuan, et al.
Published: (2024)
by: Liu, Minghuan, et al.
Published: (2024)
$\mathbb{X}$-Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity Graphs
by: Sobal, Vlad, et al.
Published: (2024)
by: Sobal, Vlad, et al.
Published: (2024)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
EgoPet: Egomotion and Interaction Data from an Animal's Perspective
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
TrajBooster: Boosting Humanoid Whole-Body Manipulation via Trajectory-Centric Learning
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning
by: Escoriza, Adrià López, et al.
Published: (2025)
by: Escoriza, Adrià López, et al.
Published: (2025)
FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions
by: Li, Peng, et al.
Published: (2026)
by: Li, Peng, et al.
Published: (2026)
Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
by: Ding, Pengxiang, et al.
Published: (2025)
by: Ding, Pengxiang, et al.
Published: (2025)
From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
by: Luo, Zhengyi, et al.
Published: (2025)
by: Luo, Zhengyi, et al.
Published: (2025)
Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations
by: Huang, Wei-Jin, et al.
Published: (2026)
by: Huang, Wei-Jin, et al.
Published: (2026)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
Visual Imitation Enables Contextual Humanoid Control
by: Allshire, Arthur, et al.
Published: (2025)
by: Allshire, Arthur, et al.
Published: (2025)
Learning and Leveraging World Models in Visual Representation Learning
by: Garrido, Quentin, et al.
Published: (2024)
by: Garrido, Quentin, et al.
Published: (2024)
Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control
by: Jia, Haozhe, et al.
Published: (2026)
by: Jia, Haozhe, et al.
Published: (2026)
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
by: Jiang, Haoran, et al.
Published: (2025)
by: Jiang, Haoran, et al.
Published: (2025)
Visually-grounded Humanoid Agents
by: Ye, Hang, et al.
Published: (2026)
by: Ye, Hang, et al.
Published: (2026)
A hierarchical loss and its problems when classifying non-hierarchically
by: Wu, Cinna, et al.
Published: (2017)
by: Wu, Cinna, et al.
Published: (2017)
Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation
by: Dong, Runpei, et al.
Published: (2026)
by: Dong, Runpei, et al.
Published: (2026)
A Recipe for Unbounded Data Augmentation in Visual Reinforcement Learning
by: Almuzairee, Abdulaziz, et al.
Published: (2024)
by: Almuzairee, Abdulaziz, et al.
Published: (2024)
Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers
by: Jauhri, Snehal, et al.
Published: (2026)
by: Jauhri, Snehal, et al.
Published: (2026)
MoDem-V2: Visuo-Motor World Models for Real-World Robot Manipulation
by: Lancaster, Patrick, et al.
Published: (2023)
by: Lancaster, Patrick, et al.
Published: (2023)
Learning Latent Action World Models In The Wild
by: Garrido, Quentin, et al.
Published: (2026)
by: Garrido, Quentin, et al.
Published: (2026)
NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
by: Liu, Fei, et al.
Published: (2025)
by: Liu, Fei, et al.
Published: (2025)
OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning
by: He, Tairan, et al.
Published: (2024)
by: He, Tairan, et al.
Published: (2024)
Scalable Trajectory Generation for Whole-Body Mobile Manipulation
by: Niu, Yida, et al.
Published: (2026)
by: Niu, Yida, et al.
Published: (2026)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
CyberDemo: Augmenting Simulated Human Demonstration for Real-World Dexterous Manipulation
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Switch-JustDance: Benchmarking Whole Body Motion Tracking Controllers Using a Commercial Console Game
by: Kim, Jeonghwan, et al.
Published: (2025)
by: Kim, Jeonghwan, et al.
Published: (2025)
The Entropy Enigma: Success and Failure of Entropy Minimization
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions
by: Xu, Sirui, et al.
Published: (2025)
by: Xu, Sirui, et al.
Published: (2025)
Similar Items
-
Learning Massively Multitask World Models for Continuous Control
by: Hansen, Nicklas, et al.
Published: (2025) -
TD-MPC2: Scalable, Robust World Models for Continuous Control
by: Hansen, Nicklas, et al.
Published: (2023) -
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025) -
Navigation World Models
by: Bar, Amir, et al.
Published: (2024) -
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)