Vidi2.5: Large Multimodal Models for Video Understanding and Creation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vidi Team, Kuo, Chia-Wen, Huang, Chuang, Du, Dawei, Chen, Fan, Lei, Fanding, Gao, Feng, Chen, Guang, Zhang, Haoji, Zhao, Haojun, Liu, Jin, Zhuge, Jingjing, Fang, Lili, Zhang, Lingxi, Wen, Longyin, Guo, Lu, Xu, Lu, Li, Lusha, Fan, Qihang, Deng, Rachel, Fang, Shaobo, Zhang, Shu, Zhu, Sijie, Siew, Stuart, Tao, Weiyan, Zhong, Wen, Shen, Xiaohui, Gu, Xin, Yuan, Ye, He, Yicheng, Cui, Yiming, Chen, Zhenfang, Wu, Zhihua, Lin, Zuhua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917210959642624
author Vidi Team
Kuo, Chia-Wen
Huang, Chuang
Du, Dawei
Chen, Fan
Lei, Fanding
Gao, Feng
Chen, Guang
Zhang, Haoji
Zhao, Haojun
Liu, Jin
Zhuge, Jingjing
Fang, Lili
Zhang, Lingxi
Wen, Longyin
Guo, Lu
Xu, Lu
Li, Lusha
Fan, Qihang
Deng, Rachel
Fang, Shaobo
Zhang, Shu
Zhu, Sijie
Siew, Stuart
Tao, Weiyan
Zhong, Wen
Shen, Xiaohui
Gu, Xin
Yuan, Ye
He, Yicheng
Cui, Yiming
Chen, Zhenfang
Wu, Zhihua
Lin, Zuhua
author_facet Vidi Team
Kuo, Chia-Wen
Huang, Chuang
Du, Dawei
Chen, Fan
Lei, Fanding
Gao, Feng
Chen, Guang
Zhang, Haoji
Zhao, Haojun
Liu, Jin
Zhuge, Jingjing
Fang, Lili
Zhang, Lingxi
Wen, Longyin
Guo, Lu
Xu, Lu
Li, Lusha
Fan, Qihang
Deng, Rachel
Fang, Shaobo
Zhang, Shu
Zhu, Sijie
Siew, Stuart
Tao, Weiyan
Zhong, Wen
Shen, Xiaohui
Gu, Xin
Yuan, Ye
He, Yicheng
Cui, Yiming
Chen, Zhenfang
Wu, Zhihua
Lin, Zuhua
contents Video has emerged as the primary medium for communication and creativity on the Internet, driving strong demand for scalable, high-quality video production. Vidi models continue to evolve toward next-generation video creation and have achieved state-of-the-art performance in multimodal temporal retrieval (TR). In its second release, Vidi2 advances video understanding with fine-grained spatio-temporal grounding (STG) and extends its capability to video question answering (Video QA), enabling comprehensive multimodal reasoning. Given a text query, Vidi2 can identify not only the corresponding timestamps but also the bounding boxes of target objects within the output time ranges. To enable comprehensive evaluation of STG, we introduce a new benchmark, VUE-STG, which offers critical improvements over existing STG datasets. In addition, we upgrade the previous VUE-TR benchmark to VUE-TR-V2, achieving a more balanced duration and query distribution. Remarkably, the Vidi2 model substantially outperforms leading proprietary systems, such as Gemini 3 Pro Preview and GPT-5, on both VUE-TR-V2 and VUE-STG, while achieving competitive results with popular open-source models with similar scale on video QA benchmarks. The latest Vidi2.5 offers significantly stronger STG capability and slightly better TR and Video QA performance over Vidi2. This update also introduces a Vidi2.5-Think model to handle plot understanding with complex plot reasoning. To comprehensively evaluate the performance of plot understanding, we propose VUE-PLOT benchmark with two tracks, Character and Reasoning. Notably, Vidi2.5-Think outperforms Gemini 3 Pro Preview on fine-grained character understanding with comparable performance on complex plot reasoning. Furthermore, we demonstrate the effectiveness of Vidi2.5 on a challenging real-world application, video editing planning.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19529
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vidi2.5: Large Multimodal Models for Video Understanding and Creation
Vidi Team
Kuo, Chia-Wen
Huang, Chuang
Du, Dawei
Chen, Fan
Lei, Fanding
Gao, Feng
Chen, Guang
Zhang, Haoji
Zhao, Haojun
Liu, Jin
Zhuge, Jingjing
Fang, Lili
Zhang, Lingxi
Wen, Longyin
Guo, Lu
Xu, Lu
Li, Lusha
Fan, Qihang
Deng, Rachel
Fang, Shaobo
Zhang, Shu
Zhu, Sijie
Siew, Stuart
Tao, Weiyan
Zhong, Wen
Shen, Xiaohui
Gu, Xin
Yuan, Ye
He, Yicheng
Cui, Yiming
Chen, Zhenfang
Wu, Zhihua
Lin, Zuhua
Computer Vision and Pattern Recognition
Video has emerged as the primary medium for communication and creativity on the Internet, driving strong demand for scalable, high-quality video production. Vidi models continue to evolve toward next-generation video creation and have achieved state-of-the-art performance in multimodal temporal retrieval (TR). In its second release, Vidi2 advances video understanding with fine-grained spatio-temporal grounding (STG) and extends its capability to video question answering (Video QA), enabling comprehensive multimodal reasoning. Given a text query, Vidi2 can identify not only the corresponding timestamps but also the bounding boxes of target objects within the output time ranges. To enable comprehensive evaluation of STG, we introduce a new benchmark, VUE-STG, which offers critical improvements over existing STG datasets. In addition, we upgrade the previous VUE-TR benchmark to VUE-TR-V2, achieving a more balanced duration and query distribution. Remarkably, the Vidi2 model substantially outperforms leading proprietary systems, such as Gemini 3 Pro Preview and GPT-5, on both VUE-TR-V2 and VUE-STG, while achieving competitive results with popular open-source models with similar scale on video QA benchmarks. The latest Vidi2.5 offers significantly stronger STG capability and slightly better TR and Video QA performance over Vidi2. This update also introduces a Vidi2.5-Think model to handle plot understanding with complex plot reasoning. To comprehensively evaluate the performance of plot understanding, we propose VUE-PLOT benchmark with two tracks, Character and Reasoning. Notably, Vidi2.5-Think outperforms Gemini 3 Pro Preview on fine-grained character understanding with comparable performance on complex plot reasoning. Furthermore, we demonstrate the effectiveness of Vidi2.5 on a challenging real-world application, video editing planning.
title Vidi2.5: Large Multimodal Models for Video Understanding and Creation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.19529