VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Deng, Andong, Du, Dawei, Chen, Zhenfang, Zhong, Wen, Chen, Fan, Chen, Guang, Kuo, Chia-Wen, Wen, Longyin, Chen, Chen, Zhu, Sijie
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913104064937984
author Deng, Andong
Du, Dawei
Chen, Zhenfang
Zhong, Wen
Chen, Fan
Chen, Guang
Kuo, Chia-Wen
Wen, Longyin
Chen, Chen
Zhu, Sijie
author_facet Deng, Andong
Du, Dawei
Chen, Zhenfang
Zhong, Wen
Chen, Fan
Chen, Guang
Kuo, Chia-Wen
Wen, Longyin
Chen, Chen
Zhu, Sijie
contents Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. While recent Large Multimodal Models (LMMs) have shown remarkable progress in general video understanding, their abilities in multi-video reasoning and operational editing workflows remain largely unexplored. We introduce VEBENCH, the first comprehensive benchmark designed to evaluate both editing knowledge understanding and operational reasoning in realistic video editing scenarios. VEBENCH contains 3.9K high-quality edited videos (over 257 hours) and 3,080 human-verified QA pairs, built through a three-round human-AI collaborative annotation pipeline that ensures precise temporal labeling and semantic consistency. It features two complementary QA tasks: 1) Video Editing Technique Recognition, assessing models' ability to identify 7 editing techniques using multimodal cues; and 2) Video Editing Operation Simulation, modeling real-world editing workflows by requiring the selection and temporal localization of relevant clips from multiple candidates. Extensive experiments across proprietary (e.g., Gemini-2.5-Pro) and open-source LMMs reveal a large gap between current model performance and human-level editing cognition. These results highlight the urgent need for bridging video understanding with creative operational reasoning. We envision VEBENCH as a foundation for advancing intelligent video editing systems and driving future research on complex reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03276
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing
Deng, Andong
Du, Dawei
Chen, Zhenfang
Zhong, Wen
Chen, Fan
Chen, Guang
Kuo, Chia-Wen
Wen, Longyin
Chen, Chen
Zhu, Sijie
Computer Vision and Pattern Recognition
Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. While recent Large Multimodal Models (LMMs) have shown remarkable progress in general video understanding, their abilities in multi-video reasoning and operational editing workflows remain largely unexplored. We introduce VEBENCH, the first comprehensive benchmark designed to evaluate both editing knowledge understanding and operational reasoning in realistic video editing scenarios. VEBENCH contains 3.9K high-quality edited videos (over 257 hours) and 3,080 human-verified QA pairs, built through a three-round human-AI collaborative annotation pipeline that ensures precise temporal labeling and semantic consistency. It features two complementary QA tasks: 1) Video Editing Technique Recognition, assessing models' ability to identify 7 editing techniques using multimodal cues; and 2) Video Editing Operation Simulation, modeling real-world editing workflows by requiring the selection and temporal localization of relevant clips from multiple candidates. Extensive experiments across proprietary (e.g., Gemini-2.5-Pro) and open-source LMMs reveal a large gap between current model performance and human-level editing cognition. These results highlight the urgent need for bridging video understanding with creative operational reasoning. We envision VEBENCH as a foundation for advancing intelligent video editing systems and driving future research on complex reasoning.
title VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.03276