UniVid: The Open-Source Unified Video Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Luo, Jiabin, Lin, Junhui, Zhang, Zeyu, Wu, Biao, Fang, Meng, Chen, Ling, Tang, Hao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912616799010816
author Luo, Jiabin
Lin, Junhui
Zhang, Zeyu
Wu, Biao
Fang, Meng
Chen, Ling
Tang, Hao
author_facet Luo, Jiabin
Lin, Junhui
Zhang, Zeyu
Wu, Biao
Fang, Meng
Chen, Ling
Tang, Hao
contents Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow-based generation due to text-visual token imbalance and the limitations of uniform cross-modal attention across the flow trajectory, and efficiently extending image-centric MLLMs to video without costly retraining. We present UniVid, a unified architecture that couples an MLLM with a diffusion decoder through a lightweight adapter, enabling both video understanding and generation. We introduce Temperature Modality Alignment to improve prompt adherence and Pyramid Reflection for efficient temporal reasoning via dynamic keyframe selection. Extensive experiments on standard benchmarks demonstrate state-of-the-art performance, achieving a 2.2% improvement on VBench-Long total score compared to EasyAnimateV5.1, and 1.0% and 3.3% accuracy gains on MSVD-QA and ActivityNet-QA, respectively, compared with the best prior 7B baselines. Code: https://github.com/AIGeeksGroup/UniVid. Website: https://aigeeksgroup.github.io/UniVid.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24200
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniVid: The Open-Source Unified Video Model
Luo, Jiabin
Lin, Junhui
Zhang, Zeyu
Wu, Biao
Fang, Meng
Chen, Ling
Tang, Hao
Computer Vision and Pattern Recognition
Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow-based generation due to text-visual token imbalance and the limitations of uniform cross-modal attention across the flow trajectory, and efficiently extending image-centric MLLMs to video without costly retraining. We present UniVid, a unified architecture that couples an MLLM with a diffusion decoder through a lightweight adapter, enabling both video understanding and generation. We introduce Temperature Modality Alignment to improve prompt adherence and Pyramid Reflection for efficient temporal reasoning via dynamic keyframe selection. Extensive experiments on standard benchmarks demonstrate state-of-the-art performance, achieving a 2.2% improvement on VBench-Long total score compared to EasyAnimateV5.1, and 1.0% and 3.3% accuracy gains on MSVD-QA and ActivityNet-QA, respectively, compared with the best prior 7B baselines. Code: https://github.com/AIGeeksGroup/UniVid. Website: https://aigeeksgroup.github.io/UniVid.
title UniVid: The Open-Source Unified Video Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.24200