LVBench: An Extreme Long Video Understanding Benchmark

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Weihan, He, Zehai, Hong, Wenyi, Cheng, Yean, Zhang, Xiaohan, Qi, Ji, Gu, Xiaotao, Huang, Shiyu, Xu, Bin, Dong, Yuxiao, Ding, Ming, Tang, Jie
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909729765195776
author Wang, Weihan
He, Zehai
Hong, Wenyi
Cheng, Yean
Zhang, Xiaohan
Qi, Ji
Gu, Xiaotao
Huang, Shiyu
Xu, Bin
Dong, Yuxiao
Ding, Ming
Tang, Jie
author_facet Wang, Weihan
He, Zehai
Hong, Wenyi
Cheng, Yean
Zhang, Xiaohan
Qi, Ji
Gu, Xiaotao
Huang, Shiyu
Xu, Bin
Dong, Yuxiao
Ding, Ming
Tang, Jie
contents Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of meeting the demands of real-world applications such as embodied intelligence for long-term decision-making, in-depth movie reviews and discussions, and live sports commentary, all of which require comprehension of long videos spanning several hours. To address this gap, we introduce LVBench, a benchmark specifically designed for long video understanding. Our dataset comprises publicly sourced videos and encompasses a diverse set of tasks aimed at long video comprehension and information extraction. LVBench is designed to challenge multimodal models to demonstrate long-term memory and extended comprehension capabilities. Our extensive evaluations reveal that current multimodal models still underperform on these demanding long video understanding tasks. Through LVBench, we aim to spur the development of more advanced models capable of tackling the complexities of long video comprehension. Our data and code are publicly available at: https://lvbench.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08035
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LVBench: An Extreme Long Video Understanding Benchmark
Wang, Weihan
He, Zehai
Hong, Wenyi
Cheng, Yean
Zhang, Xiaohan
Qi, Ji
Gu, Xiaotao
Huang, Shiyu
Xu, Bin
Dong, Yuxiao
Ding, Ming
Tang, Jie
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of meeting the demands of real-world applications such as embodied intelligence for long-term decision-making, in-depth movie reviews and discussions, and live sports commentary, all of which require comprehension of long videos spanning several hours. To address this gap, we introduce LVBench, a benchmark specifically designed for long video understanding. Our dataset comprises publicly sourced videos and encompasses a diverse set of tasks aimed at long video comprehension and information extraction. LVBench is designed to challenge multimodal models to demonstrate long-term memory and extended comprehension capabilities. Our extensive evaluations reveal that current multimodal models still underperform on these demanding long video understanding tasks. Through LVBench, we aim to spur the development of more advanced models capable of tackling the complexities of long video comprehension. Our data and code are publicly available at: https://lvbench.github.io.
title LVBench: An Extreme Long Video Understanding Benchmark
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2406.08035