Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chew, Oscar, Honcharenko, Serhii, Chen, Qian-Hui, Lu, Patricia, Zaveri, Dishant, Doan, Khoa D., Huang, Kuan-Hao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916050215370752
author Chew, Oscar
Honcharenko, Serhii
Chen, Qian-Hui
Lu, Patricia
Zaveri, Dishant
Doan, Khoa D.
Huang, Kuan-Hao
author_facet Chew, Oscar
Honcharenko, Serhii
Chen, Qian-Hui
Lu, Patricia
Zaveri, Dishant
Doan, Khoa D.
Huang, Kuan-Hao
contents A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains unclear. In this work, we introduce DistractionBench to evaluate whether VideoLLMs can robustly link subjects and events in the presence of unrelated video segments. Through controlled interventions, such as inserting short advertisement clips into longer videos, we show that VideoLLMs frequently hallucinate interactions between entities from different segments, incorrectly attributing actions from injected advertisements to subjects in the main video. We characterize this systematic hallucination as bag-of-events (BoE) behavior, where models process videos as collections of events rather than temporally structured sequences. Evaluating 11 popular VideoLLMs, we find that all models exhibit substantial BoE behavior. Our findings suggest that VideoLLMs lack reliable mechanisms for temporal grounding and motivate the development of models with more robust subject-event association.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27101
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models
Chew, Oscar
Honcharenko, Serhii
Chen, Qian-Hui
Lu, Patricia
Zaveri, Dishant
Doan, Khoa D.
Huang, Kuan-Hao
Computer Vision and Pattern Recognition
Computation and Language
A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains unclear. In this work, we introduce DistractionBench to evaluate whether VideoLLMs can robustly link subjects and events in the presence of unrelated video segments. Through controlled interventions, such as inserting short advertisement clips into longer videos, we show that VideoLLMs frequently hallucinate interactions between entities from different segments, incorrectly attributing actions from injected advertisements to subjects in the main video. We characterize this systematic hallucination as bag-of-events (BoE) behavior, where models process videos as collections of events rather than temporally structured sequences. Evaluating 11 popular VideoLLMs, we find that all models exhibit substantial BoE behavior. Our findings suggest that VideoLLMs lack reliable mechanisms for temporal grounding and motivate the development of models with more robust subject-event association.
title Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2605.27101