An Empirical Study on How Video-LLMs Answer Video Questions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gou, Chenhui, Ma, Ziyu, Duan, Zicheng, He, Haoyu, Chen, Feng, Liu, Akide, Zhuang, Bohan, Cai, Jianfei, Rezatofighi, Hamid
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912546523447296
author Gou, Chenhui
Ma, Ziyu
Duan, Zicheng
He, Haoyu
Chen, Feng
Liu, Akide
Zhuang, Bohan
Cai, Jianfei
Rezatofighi, Hamid
author_facet Gou, Chenhui
Ma, Ziyu
Duan, Zicheng
He, Haoyu
Chen, Feng
Liu, Akide
Zhuang, Bohan
Cai, Jianfei
Rezatofighi, Hamid
contents Taking advantage of large-scale data and pretrained language models, Video Large Language Models (Video-LLMs) have shown strong capabilities in answering video questions. However, most existing efforts focus on improving performance, with limited attention to understanding their internal mechanisms. This paper aims to bridge this gap through a systematic empirical study. To interpret existing VideoLLMs, we adopt attention knockouts as our primary analytical tool and design three variants: Video Temporal Knockout, Video Spatial Knockout, and Language-to-Video Knockout. Then, we apply these three knockouts on different numbers of layers (window of layers). By carefully controlling the window of layers and types of knockouts, we provide two settings: a global setting and a fine-grained setting. Our study reveals three key findings: (1) Global setting indicates Video information extraction primarily occurs in early layers, forming a clear two-stage process -- lower layers focus on perceptual encoding, while higher layers handle abstract reasoning; (2) In the fine-grained setting, certain intermediate layers exert an outsized impact on video question answering, acting as critical outliers, whereas most other layers contribute minimally; (3) In both settings, we observe that spatial-temporal modeling relies more on language-guided retrieval than on intra- and inter-frame self-attention among video tokens, despite the latter's high computational cost. Finally, we demonstrate that these insights can be leveraged to reduce attention computation in Video-LLMs. To our knowledge, this is the first work to systematically uncover how Video-LLMs internally process and understand video content, offering interpretability and efficiency perspectives for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15360
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Empirical Study on How Video-LLMs Answer Video Questions
Gou, Chenhui
Ma, Ziyu
Duan, Zicheng
He, Haoyu
Chen, Feng
Liu, Akide
Zhuang, Bohan
Cai, Jianfei
Rezatofighi, Hamid
Computer Vision and Pattern Recognition
Taking advantage of large-scale data and pretrained language models, Video Large Language Models (Video-LLMs) have shown strong capabilities in answering video questions. However, most existing efforts focus on improving performance, with limited attention to understanding their internal mechanisms. This paper aims to bridge this gap through a systematic empirical study. To interpret existing VideoLLMs, we adopt attention knockouts as our primary analytical tool and design three variants: Video Temporal Knockout, Video Spatial Knockout, and Language-to-Video Knockout. Then, we apply these three knockouts on different numbers of layers (window of layers). By carefully controlling the window of layers and types of knockouts, we provide two settings: a global setting and a fine-grained setting. Our study reveals three key findings: (1) Global setting indicates Video information extraction primarily occurs in early layers, forming a clear two-stage process -- lower layers focus on perceptual encoding, while higher layers handle abstract reasoning; (2) In the fine-grained setting, certain intermediate layers exert an outsized impact on video question answering, acting as critical outliers, whereas most other layers contribute minimally; (3) In both settings, we observe that spatial-temporal modeling relies more on language-guided retrieval than on intra- and inter-frame self-attention among video tokens, despite the latter's high computational cost. Finally, we demonstrate that these insights can be leveraged to reduce attention computation in Video-LLMs. To our knowledge, this is the first work to systematically uncover how Video-LLMs internally process and understand video content, offering interpretability and efficiency perspectives for future research.
title An Empirical Study on How Video-LLMs Answer Video Questions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.15360