Omni-Video: Democratizing Unified Video Understanding and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Zhiyu, Yang, Hao, Qin, Luozheng, Gong, Jia, Yang, Mengping, Li, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915858121490432
author Tan, Zhiyu
Yang, Hao
Qin, Luozheng
Gong, Jia
Yang, Mengping
Li, Hao
author_facet Tan, Zhiyu
Yang, Hao
Qin, Luozheng
Gong, Jia
Yang, Mengping
Li, Hao
contents Notable breakthroughs in unified understanding and generation modeling have led to remarkable advancements in image understanding, reasoning, production and editing, yet current foundational models predominantly focus on processing images, creating a gap in the development of unified models for video understanding and generation. This report presents Omni-Video, an efficient and effective unified framework for video understanding, generation, as well as instruction-based editing. Our key insight is to teach existing multimodal large language models (MLLMs) to produce continuous visual clues that are used as the input of diffusion decoders, which produce high-quality videos conditioned on these visual clues. To fully unlock the potential of our system for unified video modeling, we integrate several technical improvements: 1) a lightweight architectural design that respectively attaches a vision head on the top of MLLMs and a adapter before the input of diffusion decoders, the former produce visual tokens for the latter, which adapts these visual tokens to the conditional space of diffusion decoders; and 2) an efficient multi-stage training scheme that facilitates a fast connection between MLLMs and diffusion decoders with limited data and computational resources. We empirically demonstrate that our model exhibits satisfactory generalization abilities across video generation, editing and understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06119
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Omni-Video: Democratizing Unified Video Understanding and Generation
Tan, Zhiyu
Yang, Hao
Qin, Luozheng
Gong, Jia
Yang, Mengping
Li, Hao
Computer Vision and Pattern Recognition
Notable breakthroughs in unified understanding and generation modeling have led to remarkable advancements in image understanding, reasoning, production and editing, yet current foundational models predominantly focus on processing images, creating a gap in the development of unified models for video understanding and generation. This report presents Omni-Video, an efficient and effective unified framework for video understanding, generation, as well as instruction-based editing. Our key insight is to teach existing multimodal large language models (MLLMs) to produce continuous visual clues that are used as the input of diffusion decoders, which produce high-quality videos conditioned on these visual clues. To fully unlock the potential of our system for unified video modeling, we integrate several technical improvements: 1) a lightweight architectural design that respectively attaches a vision head on the top of MLLMs and a adapter before the input of diffusion decoders, the former produce visual tokens for the latter, which adapts these visual tokens to the conditional space of diffusion decoders; and 2) an efficient multi-stage training scheme that facilitates a fast connection between MLLMs and diffusion decoders with limited data and computational resources. We empirically demonstrate that our model exhibits satisfactory generalization abilities across video generation, editing and understanding tasks.
title Omni-Video: Democratizing Unified Video Understanding and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.06119