Discriminator-Free Direct Preference Optimization for Video Diffusion

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cheng, Haoran, Dong, Qide, Peng, Liang, Sha, Zhizhou, Feng, Weiguo, Xie, Jinghui, Song, Zhao, Wen, Shilei, He, Xiaofei, Wu, Boxi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909575612989440
author Cheng, Haoran
Dong, Qide
Peng, Liang
Sha, Zhizhou
Feng, Weiguo
Xie, Jinghui
Song, Zhao
Wen, Shilei
He, Xiaofei
Wu, Boxi
author_facet Cheng, Haoran
Dong, Qide
Peng, Liang
Sha, Zhizhou
Feng, Weiguo
Xie, Jinghui
Song, Zhao
Wen, Shilei
He, Xiaofei
Wu, Boxi
contents Direct Preference Optimization (DPO), which aligns models with human preferences through win/lose data pairs, has achieved remarkable success in language and image generation. However, applying DPO to video diffusion models faces critical challenges: (1) Data inefficiency. Generating thousands of videos per DPO iteration incurs prohibitive costs; (2) Evaluation uncertainty. Human annotations suffer from subjective bias, and automated discriminators fail to detect subtle temporal artifacts like flickering or motion incoherence. To address these, we propose a discriminator-free video DPO framework that: (1) Uses original real videos as win cases and their edited versions (e.g., reversed, shuffled, or noise-corrupted clips) as lose cases; (2) Trains video diffusion models to distinguish and avoid artifacts introduced by editing. This approach eliminates the need for costly synthetic video comparisons, provides unambiguous quality signals, and enables unlimited training data expansion through simple editing operations. We theoretically prove the framework's effectiveness even when real videos and model-generated videos follow different distributions. Experiments on CogVideoX demonstrate the efficiency of the proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08542
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Discriminator-Free Direct Preference Optimization for Video Diffusion
Cheng, Haoran
Dong, Qide
Peng, Liang
Sha, Zhizhou
Feng, Weiguo
Xie, Jinghui
Song, Zhao
Wen, Shilei
He, Xiaofei
Wu, Boxi
Computer Vision and Pattern Recognition
Direct Preference Optimization (DPO), which aligns models with human preferences through win/lose data pairs, has achieved remarkable success in language and image generation. However, applying DPO to video diffusion models faces critical challenges: (1) Data inefficiency. Generating thousands of videos per DPO iteration incurs prohibitive costs; (2) Evaluation uncertainty. Human annotations suffer from subjective bias, and automated discriminators fail to detect subtle temporal artifacts like flickering or motion incoherence. To address these, we propose a discriminator-free video DPO framework that: (1) Uses original real videos as win cases and their edited versions (e.g., reversed, shuffled, or noise-corrupted clips) as lose cases; (2) Trains video diffusion models to distinguish and avoid artifacts introduced by editing. This approach eliminates the need for costly synthetic video comparisons, provides unambiguous quality signals, and enables unlimited training data expansion through simple editing operations. We theoretically prove the framework's effectiveness even when real videos and model-generated videos follow different distributions. Experiments on CogVideoX demonstrate the efficiency of the proposed method.
title Discriminator-Free Direct Preference Optimization for Video Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.08542