LatentMove: Towards Complex Human Movement Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Taghipour, Ashkan, Ghahremani, Morteza, Bennamoun, Mohammed, Boussaid, Farid, Rekavandi, Aref Miri, Li, Zinuo, Ke, Qiuhong, Laga, Hamid
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912455562625024
author Taghipour, Ashkan
Ghahremani, Morteza
Bennamoun, Mohammed
Boussaid, Farid
Rekavandi, Aref Miri
Li, Zinuo
Ke, Qiuhong
Laga, Hamid
author_facet Taghipour, Ashkan
Ghahremani, Morteza
Bennamoun, Mohammed
Boussaid, Farid
Rekavandi, Aref Miri
Li, Zinuo
Ke, Qiuhong
Laga, Hamid
contents Image-to-video (I2V) generation seeks to produce realistic motion sequences from a single reference image. Although recent methods exhibit strong temporal consistency, they often struggle when dealing with complex, non-repetitive human movements, leading to unnatural deformations. To tackle this issue, we present LatentMove, a DiT-based framework specifically tailored for highly dynamic human animation. Our architecture incorporates a conditional control branch and learnable face/body tokens to preserve consistency as well as fine-grained details across frames. We introduce Complex-Human-Videos (CHV), a dataset featuring diverse, challenging human motions designed to benchmark the robustness of I2V systems. We also introduce two metrics to assess the flow and silhouette consistency of generated videos with their ground truth. Experimental results indicate that LatentMove substantially improves human animation quality--particularly when handling rapid, intricate movements--thereby pushing the boundaries of I2V generation. The code, the CHV dataset, and the evaluation metrics will be available at https://github.com/ --.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22046
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LatentMove: Towards Complex Human Movement Video Generation
Taghipour, Ashkan
Ghahremani, Morteza
Bennamoun, Mohammed
Boussaid, Farid
Rekavandi, Aref Miri
Li, Zinuo
Ke, Qiuhong
Laga, Hamid
Computer Vision and Pattern Recognition
Image-to-video (I2V) generation seeks to produce realistic motion sequences from a single reference image. Although recent methods exhibit strong temporal consistency, they often struggle when dealing with complex, non-repetitive human movements, leading to unnatural deformations. To tackle this issue, we present LatentMove, a DiT-based framework specifically tailored for highly dynamic human animation. Our architecture incorporates a conditional control branch and learnable face/body tokens to preserve consistency as well as fine-grained details across frames. We introduce Complex-Human-Videos (CHV), a dataset featuring diverse, challenging human motions designed to benchmark the robustness of I2V systems. We also introduce two metrics to assess the flow and silhouette consistency of generated videos with their ground truth. Experimental results indicate that LatentMove substantially improves human animation quality--particularly when handling rapid, intricate movements--thereby pushing the boundaries of I2V generation. The code, the CHV dataset, and the evaluation metrics will be available at https://github.com/ --.
title LatentMove: Towards Complex Human Movement Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22046