DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Ziyi, Kag, Anil, Skorokhodov, Ivan, Menapace, Willi, Mirzaei, Ashkan, Gilitschenski, Igor, Tulyakov, Sergey, Siarohin, Aliaksandr
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909833553248256
author Wu, Ziyi
Kag, Anil
Skorokhodov, Ivan
Menapace, Willi
Mirzaei, Ashkan
Gilitschenski, Igor
Tulyakov, Sergey
Siarohin, Aliaksandr
author_facet Wu, Ziyi
Kag, Anil
Skorokhodov, Ivan
Menapace, Willi
Mirzaei, Ashkan
Gilitschenski, Igor
Tulyakov, Sergey
Siarohin, Aliaksandr
contents Direct Preference Optimization (DPO) has recently been applied as a post-training technique for text-to-video diffusion models. To obtain training data, annotators are asked to provide preferences between two videos generated from independent noise. However, this approach prohibits fine-grained comparisons, and we point out that it biases the annotators towards low-motion clips as they often contain fewer visual artifacts. In this work, we introduce DenseDPO, a method that addresses these shortcomings by making three contributions. First, we create each video pair for DPO by denoising corrupted copies of a ground truth video. This results in aligned pairs with similar motion structures while differing in local details, effectively neutralizing the motion bias. Second, we leverage the resulting temporal alignment to label preferences on short segments rather than entire clips, yielding a denser and more precise learning signal. With only one-third of the labeled data, DenseDPO greatly improves motion generation over vanilla DPO, while matching it in text alignment, visual quality, and temporal consistency. Finally, we show that DenseDPO unlocks automatic preference annotation using off-the-shelf Vision Language Models (VLMs): GPT accurately predicts segment-level preferences similar to task-specifically fine-tuned video reward models, and DenseDPO trained on these labels achieves performance close to using human labels.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03517
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
Wu, Ziyi
Kag, Anil
Skorokhodov, Ivan
Menapace, Willi
Mirzaei, Ashkan
Gilitschenski, Igor
Tulyakov, Sergey
Siarohin, Aliaksandr
Computer Vision and Pattern Recognition
Direct Preference Optimization (DPO) has recently been applied as a post-training technique for text-to-video diffusion models. To obtain training data, annotators are asked to provide preferences between two videos generated from independent noise. However, this approach prohibits fine-grained comparisons, and we point out that it biases the annotators towards low-motion clips as they often contain fewer visual artifacts. In this work, we introduce DenseDPO, a method that addresses these shortcomings by making three contributions. First, we create each video pair for DPO by denoising corrupted copies of a ground truth video. This results in aligned pairs with similar motion structures while differing in local details, effectively neutralizing the motion bias. Second, we leverage the resulting temporal alignment to label preferences on short segments rather than entire clips, yielding a denser and more precise learning signal. With only one-third of the labeled data, DenseDPO greatly improves motion generation over vanilla DPO, while matching it in text alignment, visual quality, and temporal consistency. Finally, we show that DenseDPO unlocks automatic preference annotation using off-the-shelf Vision Language Models (VLMs): GPT accurately predicts segment-level preferences similar to task-specifically fine-tuned video reward models, and DenseDPO trained on these labels achieves performance close to using human labels.
title DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.03517