Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Jie, Gong, Xinyu, Tan, Qingshan, Li, Wen, Cheng, Yangming, Wang, Weitao, Zhan, Chenlu, Wu, Suhui, Zhang, Hao, Zhang, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909894888652800
author Du, Jie
Gong, Xinyu
Tan, Qingshan
Li, Wen
Cheng, Yangming
Wang, Weitao
Zhan, Chenlu
Wu, Suhui
Zhang, Hao
Zhang, Jun
author_facet Du, Jie
Gong, Xinyu
Tan, Qingshan
Li, Wen
Cheng, Yangming
Wang, Weitao
Zhan, Chenlu
Wu, Suhui
Zhang, Hao
Zhang, Jun
contents Recent studies have identified Direct Preference Optimization (DPO) as an efficient and reward-free approach to improving video generation quality. However, existing methods largely follow image-domain paradigms and are mainly developed on small-scale models (approximately 2B parameters), limiting their ability to address the unique challenges of video tasks, such as costly data construction, unstable training, and heavy memory consumption. To overcome these limitations, we introduce a GT-Pair that automatically builds high-quality preference pairs by using real videos as positives and model-generated videos as negatives, eliminating the need for any external annotation. We further present Reg-DPO, which incorporates the SFT loss as a regularization term into the DPO loss to enhance training stability and generation fidelity. Additionally, by combining the FSDP framework with multiple memory optimization techniques, our approach achieves nearly three times higher training capacity than using FSDP alone. Extensive experiments on both I2V and T2V tasks across multiple datasets demonstrate that our method consistently outperforms existing approaches, delivering superior video generation quality.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01450
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation
Du, Jie
Gong, Xinyu
Tan, Qingshan
Li, Wen
Cheng, Yangming
Wang, Weitao
Zhan, Chenlu
Wu, Suhui
Zhang, Hao
Zhang, Jun
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent studies have identified Direct Preference Optimization (DPO) as an efficient and reward-free approach to improving video generation quality. However, existing methods largely follow image-domain paradigms and are mainly developed on small-scale models (approximately 2B parameters), limiting their ability to address the unique challenges of video tasks, such as costly data construction, unstable training, and heavy memory consumption. To overcome these limitations, we introduce a GT-Pair that automatically builds high-quality preference pairs by using real videos as positives and model-generated videos as negatives, eliminating the need for any external annotation. We further present Reg-DPO, which incorporates the SFT loss as a regularization term into the DPO loss to enhance training stability and generation fidelity. Additionally, by combining the FSDP framework with multiple memory optimization techniques, our approach achieves nearly three times higher training capacity than using FSDP alone. Extensive experiments on both I2V and T2V tasks across multiple datasets demonstrate that our method consistently outperforms existing approaches, delivering superior video generation quality.
title Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.01450