UniVideo: Unified Understanding, Generation, and Editing for Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Cong, Liu, Quande, Ye, Zixuan, Wang, Qiulin, Wang, Xintao, Wan, Pengfei, Gai, Kun, Chen, Wenhu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909983573016576
author Wei, Cong
Liu, Quande
Ye, Zixuan
Wang, Qiulin
Wang, Xintao
Wan, Pengfei
Gai, Kun
Chen, Wenhu
author_facet Wei, Cong
Liu, Quande
Ye, Zixuan
Wang, Qiulin
Wang, Xintao
Wan, Pengfei
Gai, Kun
Chen, Wenhu
contents Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends unified modeling to the video domain. UniVideo adopts a dual-stream design, combining a Multimodal Large Language Model (MLLM) for instruction understanding with a Multimodal DiT (MMDiT) for video generation. This design preserves the MLLM's original text generation capabilities, enables accurate interpretation of complex multimodal instructions, and maintains visual consistency in the generated content. Built on this architecture, UniVideo unifies diverse video generation and editing tasks under a single multimodal instruction paradigm and is jointly trained across them. Extensive experiments demonstrate that UniVideo matches or surpasses state-of-the-art task-specific baselines in text/image-to-video generation, in-context video generation and in-context video editing. Notably, the unified design of UniVideo enables two forms of generalization. First, UniVideo supports task composition, such as combining editing with style transfer, by integrating multiple capabilities within a single instruction. Second, even without explicit training on free-form video editing, UniVideo transfers its editing capability from large-scale image editing data to this setting, handling unseen instructions such as changing the environment or altering materials within a video. Beyond these core capabilities, UniVideo also supports visual-prompt-based video generation, where the MLLM interprets visual prompts and guides the MMDiT during synthesis. To foster future research, we released our model and code.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08377
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniVideo: Unified Understanding, Generation, and Editing for Videos
Wei, Cong
Liu, Quande
Ye, Zixuan
Wang, Qiulin
Wang, Xintao
Wan, Pengfei
Gai, Kun
Chen, Wenhu
Computer Vision and Pattern Recognition
Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends unified modeling to the video domain. UniVideo adopts a dual-stream design, combining a Multimodal Large Language Model (MLLM) for instruction understanding with a Multimodal DiT (MMDiT) for video generation. This design preserves the MLLM's original text generation capabilities, enables accurate interpretation of complex multimodal instructions, and maintains visual consistency in the generated content. Built on this architecture, UniVideo unifies diverse video generation and editing tasks under a single multimodal instruction paradigm and is jointly trained across them. Extensive experiments demonstrate that UniVideo matches or surpasses state-of-the-art task-specific baselines in text/image-to-video generation, in-context video generation and in-context video editing. Notably, the unified design of UniVideo enables two forms of generalization. First, UniVideo supports task composition, such as combining editing with style transfer, by integrating multiple capabilities within a single instruction. Second, even without explicit training on free-form video editing, UniVideo transfers its editing capability from large-scale image editing data to this setting, handling unseen instructions such as changing the environment or altering materials within a video. Beyond these core capabilities, UniVideo also supports visual-prompt-based video generation, where the MLLM interprets visual prompts and guides the MMDiT during synthesis. To foster future research, we released our model and code.
title UniVideo: Unified Understanding, Generation, and Editing for Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.08377