Beyond Single Frames: Can LMMs Comprehend Temporal and Contextual Narratives in Image Sequences?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiaochen, Xia, Heming, Song, Jialin, Guan, Longyu, Yang, Yixin, Dong, Qingxiu, Luo, Weiyao, Pu, Yifan, Wang, Yiru, Meng, Xiangdi, Li, Wenjie, Sui, Zhifang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909832775204864
author Wang, Xiaochen
Xia, Heming
Song, Jialin
Guan, Longyu
Yang, Yixin
Dong, Qingxiu
Luo, Weiyao
Pu, Yifan
Wang, Yiru
Meng, Xiangdi
Li, Wenjie
Sui, Zhifang
author_facet Wang, Xiaochen
Xia, Heming
Song, Jialin
Guan, Longyu
Yang, Yixin
Dong, Qingxiu
Luo, Weiyao
Pu, Yifan
Wang, Yiru
Meng, Xiangdi
Li, Wenjie
Sui, Zhifang
contents Large Multimodal Models (LMMs) have achieved remarkable success across various visual-language tasks. However, existing benchmarks predominantly focus on single-image understanding, leaving the analysis of image sequences largely unexplored. To address this limitation, we introduce StripCipher, a comprehensive benchmark designed to evaluate capabilities of LMMs to comprehend and reason over sequential images. StripCipher comprises a human-annotated dataset and three challenging subtasks: visual narrative comprehension, contextual frame prediction, and temporal narrative reordering. Our evaluation of 16 state-of-the-art LMMs, including GPT-4o and Qwen2.5VL, reveals a significant performance gap compared to human capabilities, particularly in tasks that require reordering shuffled sequential images. For instance, GPT-4o achieves only 23.93% accuracy in the reordering subtask, which is 56.07% lower than human performance. Further quantitative analysis discuss several factors, such as input format of images, affecting the performance of LLMs in sequential understanding, underscoring the fundamental challenges that remain in the development of LMMs.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13925
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Single Frames: Can LMMs Comprehend Temporal and Contextual Narratives in Image Sequences?
Wang, Xiaochen
Xia, Heming
Song, Jialin
Guan, Longyu
Yang, Yixin
Dong, Qingxiu
Luo, Weiyao
Pu, Yifan
Wang, Yiru
Meng, Xiangdi
Li, Wenjie
Sui, Zhifang
Computation and Language
Large Multimodal Models (LMMs) have achieved remarkable success across various visual-language tasks. However, existing benchmarks predominantly focus on single-image understanding, leaving the analysis of image sequences largely unexplored. To address this limitation, we introduce StripCipher, a comprehensive benchmark designed to evaluate capabilities of LMMs to comprehend and reason over sequential images. StripCipher comprises a human-annotated dataset and three challenging subtasks: visual narrative comprehension, contextual frame prediction, and temporal narrative reordering. Our evaluation of 16 state-of-the-art LMMs, including GPT-4o and Qwen2.5VL, reveals a significant performance gap compared to human capabilities, particularly in tasks that require reordering shuffled sequential images. For instance, GPT-4o achieves only 23.93% accuracy in the reordering subtask, which is 56.07% lower than human performance. Further quantitative analysis discuss several factors, such as input format of images, affecting the performance of LLMs in sequential understanding, underscoring the fundamental challenges that remain in the development of LMMs.
title Beyond Single Frames: Can LMMs Comprehend Temporal and Contextual Narratives in Image Sequences?
topic Computation and Language
url https://arxiv.org/abs/2502.13925