Saved in:
Bibliographic Details
Main Authors: Xia, Hou, Fu, Zheren, Ling, Fangcan, Li, Jiajun, Tu, Yi, Mao, Zhendong, Zhang, Yongdong
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.19650
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916924832612352
author Xia, Hou
Fu, Zheren
Ling, Fangcan
Li, Jiajun
Tu, Yi
Mao, Zhendong
Zhang, Yongdong
author_facet Xia, Hou
Fu, Zheren
Ling, Fangcan
Li, Jiajun
Tu, Yi
Mao, Zhendong
Zhang, Yongdong
contents Large video language models (LVLMs) have made notable progress in video understanding, spurring the development of corresponding evaluation benchmarks. However, existing benchmarks generally assess overall performance across entire video sequences, overlooking nuanced behaviors such as contextual positional bias, a critical yet under-explored aspect of LVLM performance. We present Video-LevelGauge, a dedicated benchmark designed to systematically assess positional bias in LVLMs. We employ standardized probes and customized contextual setups, allowing flexible control over context length, probe position, and contextual types to simulate diverse real-world scenarios. In addition, we introduce a comprehensive analysis method that combines statistical measures with morphological pattern recognition to characterize bias. Our benchmark comprises 438 manually curated videos spanning multiple types, yielding 1,177 high-quality multiple-choice questions and 120 open-ended questions, validated for their effectiveness in exposing positional bias. Based on these, we evaluate 27 state-of-the-art LVLMs, including both commercial and open-source models. Our findings reveal significant positional biases in many leading open-source models, typically exhibiting head or neighbor-content preferences. In contrast, commercial models such as Gemini2.5-Pro show impressive, consistent performance across entire video sequences. Further analyses on context length, context variation, and model scale provide actionable insights for mitigating bias and guiding model enhancement . https://github.com/Cola-any/Video-LevelGauge
format Preprint
id arxiv_https___arxiv_org_abs_2508_19650
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
Xia, Hou
Fu, Zheren
Ling, Fangcan
Li, Jiajun
Tu, Yi
Mao, Zhendong
Zhang, Yongdong
Computer Vision and Pattern Recognition
Large video language models (LVLMs) have made notable progress in video understanding, spurring the development of corresponding evaluation benchmarks. However, existing benchmarks generally assess overall performance across entire video sequences, overlooking nuanced behaviors such as contextual positional bias, a critical yet under-explored aspect of LVLM performance. We present Video-LevelGauge, a dedicated benchmark designed to systematically assess positional bias in LVLMs. We employ standardized probes and customized contextual setups, allowing flexible control over context length, probe position, and contextual types to simulate diverse real-world scenarios. In addition, we introduce a comprehensive analysis method that combines statistical measures with morphological pattern recognition to characterize bias. Our benchmark comprises 438 manually curated videos spanning multiple types, yielding 1,177 high-quality multiple-choice questions and 120 open-ended questions, validated for their effectiveness in exposing positional bias. Based on these, we evaluate 27 state-of-the-art LVLMs, including both commercial and open-source models. Our findings reveal significant positional biases in many leading open-source models, typically exhibiting head or neighbor-content preferences. In contrast, commercial models such as Gemini2.5-Pro show impressive, consistent performance across entire video sequences. Further analyses on context length, context variation, and model scale provide actionable insights for mitigating bias and guiding model enhancement . https://github.com/Cola-any/Video-LevelGauge
title Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.19650