DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Wonjoon, Won, Jiyun, Han, Janghyeok, Dai, Qi, Luo, Chong, Baek, Seung-Hwan, Cho, Sunghyun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912998002524160
author Jin, Wonjoon
Won, Jiyun
Han, Janghyeok
Dai, Qi
Luo, Chong
Baek, Seung-Hwan
Cho, Sunghyun
author_facet Jin, Wonjoon
Won, Jiyun
Han, Janghyeok
Dai, Qi
Luo, Chong
Baek, Seung-Hwan
Cho, Sunghyun
contents Despite recent progress, video diffusion models still struggle to synthesize realistic videos involving highly dynamic motions or requiring fine-grained motion controllability. A central limitation lies in the scarcity of such examples in commonly used training datasets. To address this, we introduce DynaVid, a video synthesis framework that leverages synthetic motion data in training, which is represented as optical flow and rendered using computer graphics pipelines. This approach offers two key advantages. First, synthetic motion offers diverse motion patterns and precise control signals that are difficult to obtain from real data. Second, unlike rendered videos with artificial appearances, rendered optical flow encodes only motion and is decoupled from appearance, thereby preventing models from reproducing the unnatural look of synthetic videos. Building on this idea, DynaVid adopts a two-stage generation framework: a motion generator first synthesizes motion, and then a motion-guided video generator produces video frames conditioned on that motion. This decoupled formulation enables the model to learn dynamic motion patterns from synthetic data while preserving visual realism from real-world videos. We validate our framework on two challenging scenarios, vigorous human motion generation and extreme camera motion control, where existing datasets are particularly limited. Extensive experiments demonstrate that DynaVid improves the realism and controllability in dynamic motion generation and camera motion control.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01666
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data
Jin, Wonjoon
Won, Jiyun
Han, Janghyeok
Dai, Qi
Luo, Chong
Baek, Seung-Hwan
Cho, Sunghyun
Computer Vision and Pattern Recognition
Despite recent progress, video diffusion models still struggle to synthesize realistic videos involving highly dynamic motions or requiring fine-grained motion controllability. A central limitation lies in the scarcity of such examples in commonly used training datasets. To address this, we introduce DynaVid, a video synthesis framework that leverages synthetic motion data in training, which is represented as optical flow and rendered using computer graphics pipelines. This approach offers two key advantages. First, synthetic motion offers diverse motion patterns and precise control signals that are difficult to obtain from real data. Second, unlike rendered videos with artificial appearances, rendered optical flow encodes only motion and is decoupled from appearance, thereby preventing models from reproducing the unnatural look of synthetic videos. Building on this idea, DynaVid adopts a two-stage generation framework: a motion generator first synthesizes motion, and then a motion-guided video generator produces video frames conditioned on that motion. This decoupled formulation enables the model to learn dynamic motion patterns from synthetic data while preserving visual realism from real-world videos. We validate our framework on two challenging scenarios, vigorous human motion generation and extreme camera motion control, where existing datasets are particularly limited. Extensive experiments demonstrate that DynaVid improves the realism and controllability in dynamic motion generation and camera motion control.
title DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.01666