Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Yatai, Zhang, Jiacheng, Wu, Jie, Zhang, Shilong, Chen, Shoufa, GE, Chongjian, Sun, Peize, Chen, Weifeng, Shao, Wenqi, Xiao, Xuefeng, Huang, Weilin, Luo, Ping
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915071661178880
author Ji, Yatai
Zhang, Jiacheng
Wu, Jie
Zhang, Shilong
Chen, Shoufa
GE, Chongjian
Sun, Peize
Chen, Weifeng
Shao, Wenqi
Xiao, Xuefeng
Huang, Weilin
Luo, Ping
author_facet Ji, Yatai
Zhang, Jiacheng
Wu, Jie
Zhang, Shilong
Chen, Shoufa
GE, Chongjian
Sun, Peize
Chen, Weifeng
Shao, Wenqi
Xiao, Xuefeng
Huang, Weilin
Luo, Ping
contents Text-to-video models have made remarkable advancements through optimization on high-quality text-video pairs, where the textual prompts play a pivotal role in determining quality of output videos. However, achieving the desired output often entails multiple revisions and iterative inference to refine user-provided prompts. Current automatic methods for refining prompts encounter challenges such as Modality-Inconsistency, Cost-Discrepancy, and Model-Unaware when applied to text-to-video diffusion models. To address these problem, we introduce an LLM-based prompt adaptation framework, termed as Prompt-A-Video, which excels in crafting Video-Centric, Labor-Free and Preference-Aligned prompts tailored to specific video diffusion model. Our approach involves a meticulously crafted two-stage optimization and alignment system. Initially, we conduct a reward-guided prompt evolution pipeline to automatically create optimal prompts pool and leverage them for supervised fine-tuning (SFT) of the LLM. Then multi-dimensional rewards are employed to generate pairwise data for the SFT model, followed by the direct preference optimization (DPO) algorithm to further facilitate preference alignment. Through extensive experimentation and comparative analyses, we validate the effectiveness of Prompt-A-Video across diverse generation models, highlighting its potential to push the boundaries of video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15156
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
Ji, Yatai
Zhang, Jiacheng
Wu, Jie
Zhang, Shilong
Chen, Shoufa
GE, Chongjian
Sun, Peize
Chen, Weifeng
Shao, Wenqi
Xiao, Xuefeng
Huang, Weilin
Luo, Ping
Computer Vision and Pattern Recognition
Computation and Language
Multimedia
Text-to-video models have made remarkable advancements through optimization on high-quality text-video pairs, where the textual prompts play a pivotal role in determining quality of output videos. However, achieving the desired output often entails multiple revisions and iterative inference to refine user-provided prompts. Current automatic methods for refining prompts encounter challenges such as Modality-Inconsistency, Cost-Discrepancy, and Model-Unaware when applied to text-to-video diffusion models. To address these problem, we introduce an LLM-based prompt adaptation framework, termed as Prompt-A-Video, which excels in crafting Video-Centric, Labor-Free and Preference-Aligned prompts tailored to specific video diffusion model. Our approach involves a meticulously crafted two-stage optimization and alignment system. Initially, we conduct a reward-guided prompt evolution pipeline to automatically create optimal prompts pool and leverage them for supervised fine-tuning (SFT) of the LLM. Then multi-dimensional rewards are employed to generate pairwise data for the SFT model, followed by the direct preference optimization (DPO) algorithm to further facilitate preference alignment. Through extensive experimentation and comparative analyses, we validate the effectiveness of Prompt-A-Video across diverse generation models, highlighting its potential to push the boundaries of video generation.
title Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
topic Computer Vision and Pattern Recognition
Computation and Language
Multimedia
url https://arxiv.org/abs/2412.15156