Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wu, Shengqiong, Ye, Weicai, Wang, Jiahao, Liu, Quande, Wang, Xintao, Wan, Pengfei, Zhang, Di, Gai, Kun, Yan, Shuicheng, Fei, Hao, Chua, Tat-Seng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915670119153664
author Wu, Shengqiong
Ye, Weicai
Wang, Jiahao
Liu, Quande
Wang, Xintao
Wan, Pengfei
Zhang, Di
Gai, Kun
Yan, Shuicheng
Fei, Hao
Chua, Tat-Seng
author_facet Wu, Shengqiong
Ye, Weicai
Wang, Jiahao
Liu, Quande
Wang, Xintao
Wan, Pengfei
Zhang, Di
Gai, Kun
Yan, Shuicheng
Fei, Hao
Chua, Tat-Seng
contents To address the bottleneck of accurate user intent interpretation within the current video generation community, we present Any2Caption, a novel framework for controllable video generation under any condition. The key idea is to decouple various condition interpretation steps from the video synthesis step. By leveraging modern multimodal large language models (MLLMs), Any2Caption interprets diverse inputs--text, images, videos, and specialized cues such as region, motion, and camera poses--into dense, structured captions that offer backbone video generators with better guidance. We also introduce Any2CapIns, a large-scale dataset with 337K instances and 407K conditions for any-condition-to-caption instruction tuning. Comprehensive evaluations demonstrate significant improvements of our system in controllability and video quality across various aspects of existing video generation models. Project Page: https://sqwu.top/Any2Cap/
format Preprint
id arxiv_https___arxiv_org_abs_2503_24379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
Wu, Shengqiong
Ye, Weicai
Wang, Jiahao
Liu, Quande
Wang, Xintao
Wan, Pengfei
Zhang, Di
Gai, Kun
Yan, Shuicheng
Fei, Hao
Chua, Tat-Seng
Computer Vision and Pattern Recognition
Artificial Intelligence
To address the bottleneck of accurate user intent interpretation within the current video generation community, we present Any2Caption, a novel framework for controllable video generation under any condition. The key idea is to decouple various condition interpretation steps from the video synthesis step. By leveraging modern multimodal large language models (MLLMs), Any2Caption interprets diverse inputs--text, images, videos, and specialized cues such as region, motion, and camera poses--into dense, structured captions that offer backbone video generators with better guidance. We also introduce Any2CapIns, a large-scale dataset with 337K instances and 407K conditions for any-condition-to-caption instruction tuning. Comprehensive evaluations demonstrate significant improvements of our system in controllability and video quality across various aspects of existing video generation models. Project Page: https://sqwu.top/Any2Cap/
title Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.24379