HINT: Hierarchical Interaction Modeling for Autoregressive Multi-Human Motion Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Mengge, Di, Yan, Wang, Gu, Qu, Yun, Zhu, Dekai, Li, Yanyan, Ji, Xiangyang
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914286142488576
author Liu, Mengge
Di, Yan
Wang, Gu
Qu, Yun
Zhu, Dekai
Li, Yanyan
Ji, Xiangyang
author_facet Liu, Mengge
Di, Yan
Wang, Gu
Qu, Yun
Zhu, Dekai
Li, Yanyan
Ji, Xiangyang
contents Text-driven multi-human motion generation with complex interactions remains a challenging problem. Despite progress in performance, existing offline methods that generate fixed-length motions with a fixed number of agents, are inherently limited in handling long or variable text, and varying agent counts. These limitations naturally encourage autoregressive formulations, which predict future motions step by step conditioned on all past trajectories and current text guidance. In this work, we introduce HINT, the first autoregressive framework for multi-human motion generation with Hierarchical INTeraction modeling in diffusion. First, HINT leverages a disentangled motion representation within a canonicalized latent space, decoupling local motion semantics from inter-person interactions. This design facilitates direct adaptation to varying numbers of human participants without requiring additional refinement. Second, HINT adopts a sliding-window strategy for efficient online generation, and aggregates local within-window and global cross-window conditions to capture past human history, inter-person dependencies, and align with text guidance. This strategy not only enables fine-grained interaction modeling within each window but also preserves long-horizon coherence across all the long sequence. Extensive experiments on public benchmarks demonstrate that HINT matches the performance of strong offline models and surpasses autoregressive baselines. Notably, on InterHuman, HINT achieves an FID of 3.100, significantly improving over the previous state-of-the-art score of 5.154.
format Preprint
id arxiv_https___arxiv_org_abs_2601_20383
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HINT: Hierarchical Interaction Modeling for Autoregressive Multi-Human Motion Generation
Liu, Mengge
Di, Yan
Wang, Gu
Qu, Yun
Zhu, Dekai
Li, Yanyan
Ji, Xiangyang
Computer Vision and Pattern Recognition
Text-driven multi-human motion generation with complex interactions remains a challenging problem. Despite progress in performance, existing offline methods that generate fixed-length motions with a fixed number of agents, are inherently limited in handling long or variable text, and varying agent counts. These limitations naturally encourage autoregressive formulations, which predict future motions step by step conditioned on all past trajectories and current text guidance. In this work, we introduce HINT, the first autoregressive framework for multi-human motion generation with Hierarchical INTeraction modeling in diffusion. First, HINT leverages a disentangled motion representation within a canonicalized latent space, decoupling local motion semantics from inter-person interactions. This design facilitates direct adaptation to varying numbers of human participants without requiring additional refinement. Second, HINT adopts a sliding-window strategy for efficient online generation, and aggregates local within-window and global cross-window conditions to capture past human history, inter-person dependencies, and align with text guidance. This strategy not only enables fine-grained interaction modeling within each window but also preserves long-horizon coherence across all the long sequence. Extensive experiments on public benchmarks demonstrate that HINT matches the performance of strong offline models and surpasses autoregressive baselines. Notably, on InterHuman, HINT achieves an FID of 3.100, significantly improving over the previous state-of-the-art score of 5.154.
title HINT: Hierarchical Interaction Modeling for Autoregressive Multi-Human Motion Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.20383