RDPO: Real Data Preference Optimization for Physics Consistency Video Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Qian, Wenxu, Wang, Chaoyue, Peng, Hou, Tan, Zhiyu, Li, Hao, Zeng, Anxiang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909657394577408
author Qian, Wenxu
Wang, Chaoyue
Peng, Hou
Tan, Zhiyu
Li, Hao
Zeng, Anxiang
author_facet Qian, Wenxu
Wang, Chaoyue
Peng, Hou
Tan, Zhiyu
Li, Hao
Zeng, Anxiang
contents Video generation techniques have achieved remarkable advancements in visual quality, yet faithfully reproducing real-world physics remains elusive. Preference-based model post-training may improve physical consistency, but requires costly human-annotated datasets or reward models that are not yet feasible. To address these challenges, we present Real Data Preference Optimisation (RDPO), an annotation-free framework that distills physical priors directly from real-world videos. Specifically, the proposed RDPO reverse-samples real video sequences with a pre-trained generator to automatically build preference pairs that are statistically distinguishable in terms of physical correctness. A multi-stage iterative training schedule then guides the generator to obey physical laws increasingly well. Benefiting from the dynamic information explored from real videos, our proposed RDPO significantly improves the action coherence and physical realism of the generated videos. Evaluations on multiple benchmarks and human evaluations have demonstrated that RDPO achieves improvements across multiple dimensions. The source code and demonstration of this paper are available at: https://wwenxu.github.io/RDPO/
format Preprint
id arxiv_https___arxiv_org_abs_2506_18655
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
Qian, Wenxu
Wang, Chaoyue
Peng, Hou
Tan, Zhiyu
Li, Hao
Zeng, Anxiang
Computer Vision and Pattern Recognition
I.2.6; I.2.10
Video generation techniques have achieved remarkable advancements in visual quality, yet faithfully reproducing real-world physics remains elusive. Preference-based model post-training may improve physical consistency, but requires costly human-annotated datasets or reward models that are not yet feasible. To address these challenges, we present Real Data Preference Optimisation (RDPO), an annotation-free framework that distills physical priors directly from real-world videos. Specifically, the proposed RDPO reverse-samples real video sequences with a pre-trained generator to automatically build preference pairs that are statistically distinguishable in terms of physical correctness. A multi-stage iterative training schedule then guides the generator to obey physical laws increasingly well. Benefiting from the dynamic information explored from real videos, our proposed RDPO significantly improves the action coherence and physical realism of the generated videos. Evaluations on multiple benchmarks and human evaluations have demonstrated that RDPO achieves improvements across multiple dimensions. The source code and demonstration of this paper are available at: https://wwenxu.github.io/RDPO/
title RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
topic Computer Vision and Pattern Recognition
I.2.6; I.2.10
url https://arxiv.org/abs/2506.18655