SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zhenyu, Hu, Yuhang, Du, Zemin, Xue, Dizhan, Qian, Shengsheng, Wu, Jiahong, Yang, Fan, Dong, Weiming, Xu, Changsheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911267213541376
author Yang, Zhenyu
Hu, Yuhang
Du, Zemin
Xue, Dizhan
Qian, Shengsheng
Wu, Jiahong
Yang, Fan
Dong, Weiming
Xu, Changsheng
author_facet Yang, Zhenyu
Hu, Yuhang
Du, Zemin
Xue, Dizhan
Qian, Shengsheng
Wu, Jiahong
Yang, Fan
Dong, Weiming
Xu, Changsheng
contents Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-context streaming video understanding. Current benchmarks for video understanding typically emphasize isolated single-instance text inputs and fail to evaluate the capacity to sustain temporal reasoning throughout the entire duration of video streams. To address these limitations, we introduce SVBench, a pioneering benchmark with temporal multi-turn question-answering chains specifically designed to thoroughly assess the capabilities of streaming video understanding of current LVLMs. We design a semi-automated annotation pipeline to obtain 49,979 Question-Answer (QA) pairs of 1,353 streaming videos, which includes generating QA chains that represent a series of consecutive multi-turn dialogues over video segments and constructing temporal linkages between successive QA chains. Our experimental results, obtained from 14 models in dialogue and streaming evaluations, reveal that while the closed-source GPT-4o outperforms others, most open-source LVLMs struggle with long-context streaming video understanding. We also construct a StreamingChat model, which significantly outperforms open-source LVLMs on our SVBench and achieves comparable performance on diverse vision-language benchmarks. We expect SVBench to advance the research of streaming video understanding by providing a comprehensive and in-depth analysis of current LVLMs. Our benchmark and model can be accessed at https://github.com/sotayang/SVBench.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10810
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
Yang, Zhenyu
Hu, Yuhang
Du, Zemin
Xue, Dizhan
Qian, Shengsheng
Wu, Jiahong
Yang, Fan
Dong, Weiming
Xu, Changsheng
Computer Vision and Pattern Recognition
Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-context streaming video understanding. Current benchmarks for video understanding typically emphasize isolated single-instance text inputs and fail to evaluate the capacity to sustain temporal reasoning throughout the entire duration of video streams. To address these limitations, we introduce SVBench, a pioneering benchmark with temporal multi-turn question-answering chains specifically designed to thoroughly assess the capabilities of streaming video understanding of current LVLMs. We design a semi-automated annotation pipeline to obtain 49,979 Question-Answer (QA) pairs of 1,353 streaming videos, which includes generating QA chains that represent a series of consecutive multi-turn dialogues over video segments and constructing temporal linkages between successive QA chains. Our experimental results, obtained from 14 models in dialogue and streaming evaluations, reveal that while the closed-source GPT-4o outperforms others, most open-source LVLMs struggle with long-context streaming video understanding. We also construct a StreamingChat model, which significantly outperforms open-source LVLMs on our SVBench and achieves comparable performance on diverse vision-language benchmarks. We expect SVBench to advance the research of streaming video understanding by providing a comprehensive and in-depth analysis of current LVLMs. Our benchmark and model can be accessed at https://github.com/sotayang/SVBench.
title SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.10810