VideoDPO: Omni-Preference Alignment for Video Diffusion Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Runtao, Wu, Haoyu, Ziqiang, Zheng, Wei, Chen, He, Yingqing, Pi, Renjie, Chen, Qifeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916531184599040
author Liu, Runtao
Wu, Haoyu
Ziqiang, Zheng
Wei, Chen
He, Yingqing
Pi, Renjie
Chen, Qifeng
author_facet Liu, Runtao
Wu, Haoyu
Ziqiang, Zheng
Wei, Chen
He, Yingqing
Pi, Renjie
Chen, Qifeng
contents Recent progress in generative diffusion models has greatly advanced text-to-video generation. While text-to-video models trained on large-scale, diverse datasets can produce varied outputs, these generations often deviate from user preferences, highlighting the need for preference alignment on pre-trained models. Although Direct Preference Optimization (DPO) has demonstrated significant improvements in language and image generation, we pioneer its adaptation to video diffusion models and propose a VideoDPO pipeline by making several key adjustments. Unlike previous image alignment methods that focus solely on either (i) visual quality or (ii) semantic alignment between text and videos, we comprehensively consider both dimensions and construct a preference score accordingly, which we term the OmniScore. We design a pipeline to automatically collect preference pair data based on the proposed OmniScore and discover that re-weighting these pairs based on the score significantly impacts overall preference alignment. Our experiments demonstrate substantial improvements in both visual quality and semantic alignment, ensuring that no preference aspect is neglected. Code and data will be shared at https://videodpo.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14167
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
Liu, Runtao
Wu, Haoyu
Ziqiang, Zheng
Wei, Chen
He, Yingqing
Pi, Renjie
Chen, Qifeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recent progress in generative diffusion models has greatly advanced text-to-video generation. While text-to-video models trained on large-scale, diverse datasets can produce varied outputs, these generations often deviate from user preferences, highlighting the need for preference alignment on pre-trained models. Although Direct Preference Optimization (DPO) has demonstrated significant improvements in language and image generation, we pioneer its adaptation to video diffusion models and propose a VideoDPO pipeline by making several key adjustments. Unlike previous image alignment methods that focus solely on either (i) visual quality or (ii) semantic alignment between text and videos, we comprehensively consider both dimensions and construct a preference score accordingly, which we term the OmniScore. We design a pipeline to automatically collect preference pair data based on the proposed OmniScore and discover that re-weighting these pairs based on the score significantly impacts overall preference alignment. Our experiments demonstrate substantial improvements in both visual quality and semantic alignment, ensuring that no preference aspect is neglected. Code and data will be shared at https://videodpo.github.io/.
title VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.14167