Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Taghipour, Ashkan, Ghahremani, Morteza, Li, Zinuo, Laga, Hamid, Boussaid, Farid, Bennamoun, Mohammed
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910046164615168
author Taghipour, Ashkan
Ghahremani, Morteza
Li, Zinuo
Laga, Hamid
Boussaid, Farid
Bennamoun, Mohammed
author_facet Taghipour, Ashkan
Ghahremani, Morteza
Li, Zinuo
Laga, Hamid
Boussaid, Farid
Bennamoun, Mohammed
contents Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit pose-based controls, though effective, require users to provide complete skeleton sequences that are costly to produce for long and dynamic actions. We propose a two-stage cascaded framework that addresses both limitations. First, an autoregressive text-to-skeleton model generates 2D pose sequences from natural language descriptions by predicting each joint conditioned on previously generated poses. This design captures long-range temporal dependencies and inter-joint coordination required for complex motions. Second, a pose-conditioned video diffusion model synthesizes videos from a reference image and the generated skeleton sequence. It employs DINO-ALF (Adaptive Layer Fusion), a multi-level reference encoder that preserves appearance and clothing details under large pose changes and self-occlusions. To address the lack of publicly available datasets for complex human motion video generation, we introduce a Blender-based synthetic dataset containing 2,000 videos with diverse characters performing acrobatic and stunt-like motions. The dataset provides full control over appearance, motion, and environment. It fills an important gap because existing benchmarks significantly under-represent acrobatic motions while web-collected datasets raise copyright and privacy concerns. Experiments on our synthetic dataset and the Motion-X Fitness benchmark show that our text-to-skeleton model outperforms prior methods on FID, R-precision, and motion diversity. Our pose-to-video model also achieves the best results among all compared methods on VBench metrics for temporal consistency, motion smoothness, and subject preservation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08028
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
Taghipour, Ashkan
Ghahremani, Morteza
Li, Zinuo
Laga, Hamid
Boussaid, Farid
Bennamoun, Mohammed
Computer Vision and Pattern Recognition
Multimedia
Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit pose-based controls, though effective, require users to provide complete skeleton sequences that are costly to produce for long and dynamic actions. We propose a two-stage cascaded framework that addresses both limitations. First, an autoregressive text-to-skeleton model generates 2D pose sequences from natural language descriptions by predicting each joint conditioned on previously generated poses. This design captures long-range temporal dependencies and inter-joint coordination required for complex motions. Second, a pose-conditioned video diffusion model synthesizes videos from a reference image and the generated skeleton sequence. It employs DINO-ALF (Adaptive Layer Fusion), a multi-level reference encoder that preserves appearance and clothing details under large pose changes and self-occlusions. To address the lack of publicly available datasets for complex human motion video generation, we introduce a Blender-based synthetic dataset containing 2,000 videos with diverse characters performing acrobatic and stunt-like motions. The dataset provides full control over appearance, motion, and environment. It fills an important gap because existing benchmarks significantly under-represent acrobatic motions while web-collected datasets raise copyright and privacy concerns. Experiments on our synthetic dataset and the Motion-X Fitness benchmark show that our text-to-skeleton model outperforms prior methods on FID, R-precision, and motion diversity. Our pose-to-video model also achieves the best results among all compared methods on VBench metrics for temporal consistency, motion smoothness, and subject preservation.
title Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2603.08028