Generating Human Motion Videos using a Cascaded Text-to-Video Framework

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nam, Hyelin, Go, Hyojun, Park, Byeongjun, Kim, Byung-Hoon, Chung, Hyungjin
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908577098104832
author Nam, Hyelin
Go, Hyojun
Park, Byeongjun
Kim, Byung-Hoon
Chung, Hyungjin
author_facet Nam, Hyelin
Go, Hyojun
Park, Byeongjun
Kim, Byung-Hoon
Chung, Hyungjin
contents Human video generation is becoming an increasingly important task with broad applications in graphics, entertainment, and embodied AI. Despite the rapid progress of video diffusion models (VDMs), their use for general-purpose human video generation remains underexplored, with most works constrained to image-to-video setups or narrow domains like dance videos. In this work, we propose CAMEO, a cascaded framework for general human motion video generation. It seamlessly bridges Text-to-Motion (T2M) models and conditional VDMs, mitigating suboptimal factors that may arise in this process across both training and inference through carefully designed components. Specifically, we analyze and prepare both textual prompts and visual conditions to effectively train the VDM, ensuring robust alignment between motion descriptions, conditioning signals, and the generated videos. Furthermore, we introduce a camera-aware conditioning module that connects the two stages, automatically selecting viewpoints aligned with the input text to enhance coherence and reduce manual intervention. We demonstrate the effectiveness of our approach on both the MovieGen benchmark and a newly introduced benchmark tailored to the T2M-VDM combination, while highlighting its versatility across diverse use cases.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03909
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generating Human Motion Videos using a Cascaded Text-to-Video Framework
Nam, Hyelin
Go, Hyojun
Park, Byeongjun
Kim, Byung-Hoon
Chung, Hyungjin
Computer Vision and Pattern Recognition
Human video generation is becoming an increasingly important task with broad applications in graphics, entertainment, and embodied AI. Despite the rapid progress of video diffusion models (VDMs), their use for general-purpose human video generation remains underexplored, with most works constrained to image-to-video setups or narrow domains like dance videos. In this work, we propose CAMEO, a cascaded framework for general human motion video generation. It seamlessly bridges Text-to-Motion (T2M) models and conditional VDMs, mitigating suboptimal factors that may arise in this process across both training and inference through carefully designed components. Specifically, we analyze and prepare both textual prompts and visual conditions to effectively train the VDM, ensuring robust alignment between motion descriptions, conditioning signals, and the generated videos. Furthermore, we introduce a camera-aware conditioning module that connects the two stages, automatically selecting viewpoints aligned with the input text to enhance coherence and reduce manual intervention. We demonstrate the effectiveness of our approach on both the MovieGen benchmark and a newly introduced benchmark tailored to the T2M-VDM combination, while highlighting its versatility across diverse use cases.
title Generating Human Motion Videos using a Cascaded Text-to-Video Framework
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.03909