MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Bin, Zhang, Ruichi, Liang, Han, Zhang, Jingyan, Zhang, Juze, Chen, Xin, Wang, Jingya
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910255578873856
author Li, Bin
Zhang, Ruichi
Liang, Han
Zhang, Jingyan
Zhang, Juze
Chen, Xin
Wang, Jingya
author_facet Li, Bin
Zhang, Ruichi
Liang, Han
Zhang, Jingyan
Zhang, Juze
Chen, Xin
Wang, Jingya
contents Enabling physics-based humanoids to execute diverse behaviors from high-level textual commands remains a significant challenge. Existing methods typically follow either a two-stage paradigm that combines kinematic motion generation with physics-based tracking, or an end-to-end imitation-learning paradigm that directly generates actions from text. However, the former suffers from the inherent domain shift between kinematic generation and physics-based tracking, while the latter struggles with the substantial modality gap between textual commands and low-level actions, limiting effective semantic alignment. Notably, humanoid states encode rich motion dynamics that are more semantically aligned with textual descriptions than low-level actions, making them a natural basis for deriving behavioral intent. Building upon this insight, we propose MIND, a novel end-to-end diffusion framework for text-driven physics-based humanoid control that leverages behavioral intent as a semantic bridge between textual commands and low-level actions. At its core, MIND introduces a multi-scale intent diffusion mechanism, where a holistic intent predictor captures global behavioral dynamics to guide overall behavior synthesis, while an immediate intent predictor provides step-wise, fine-grained signals for local behavior refinement at each diffusion step. This hierarchical intent formulation imposes a structured inductive bias for humanoid control, improving semantic alignment and behavioral naturalness. Furthermore, MIND encodes humanoid states into a latent space to enable more effective semantic intent modeling. Extensive experiments demonstrate that MIND outperforms existing methods and synthesizes coherent, physically plausible, and semantically aligned humanoid behaviors from text commands. Our code will be released to facilitate future research.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26006
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control
Li, Bin
Zhang, Ruichi
Liang, Han
Zhang, Jingyan
Zhang, Juze
Chen, Xin
Wang, Jingya
Computer Vision and Pattern Recognition
Graphics
Robotics
Enabling physics-based humanoids to execute diverse behaviors from high-level textual commands remains a significant challenge. Existing methods typically follow either a two-stage paradigm that combines kinematic motion generation with physics-based tracking, or an end-to-end imitation-learning paradigm that directly generates actions from text. However, the former suffers from the inherent domain shift between kinematic generation and physics-based tracking, while the latter struggles with the substantial modality gap between textual commands and low-level actions, limiting effective semantic alignment. Notably, humanoid states encode rich motion dynamics that are more semantically aligned with textual descriptions than low-level actions, making them a natural basis for deriving behavioral intent. Building upon this insight, we propose MIND, a novel end-to-end diffusion framework for text-driven physics-based humanoid control that leverages behavioral intent as a semantic bridge between textual commands and low-level actions. At its core, MIND introduces a multi-scale intent diffusion mechanism, where a holistic intent predictor captures global behavioral dynamics to guide overall behavior synthesis, while an immediate intent predictor provides step-wise, fine-grained signals for local behavior refinement at each diffusion step. This hierarchical intent formulation imposes a structured inductive bias for humanoid control, improving semantic alignment and behavioral naturalness. Furthermore, MIND encodes humanoid states into a latent space to enable more effective semantic intent modeling. Extensive experiments demonstrate that MIND outperforms existing methods and synthesizes coherent, physically plausible, and semantically aligned humanoid behaviors from text commands. Our code will be released to facilitate future research.
title MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control
topic Computer Vision and Pattern Recognition
Graphics
Robotics
url https://arxiv.org/abs/2605.26006