VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Hongbo, Wang, Meng, Zhu, Fei, Liu, Wenzhuo, Ni, Bolin, Zeng, Fanhu, Meng, Gaofeng, Zhang, Zhaoxiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914216642871296
author Zhao, Hongbo
Wang, Meng
Zhu, Fei
Liu, Wenzhuo
Ni, Bolin
Zeng, Fanhu
Meng, Gaofeng
Zhang, Zhaoxiang
author_facet Zhao, Hongbo
Wang, Meng
Zhu, Fei
Liu, Wenzhuo
Ni, Bolin
Zeng, Fanhu
Meng, Gaofeng
Zhang, Zhaoxiang
contents The computational and memory overheads associated with expanding the context window of LLMs severely limit their scalability. A noteworthy solution is vision-text compression (VTC), exemplified by frameworks like DeepSeek-OCR and Glyph, which convert long texts into dense 2D visual representations, thereby achieving token compression ratios of 3x-20x. However, the impact of this high information density on the core long-context capabilities of vision-language models (VLMs) remains under-investigated. To address this gap, we introduce the first benchmark for VTC and systematically assess the performance of VLMs across three long-context understanding settings: VTC-Retrieval, which evaluates the model's ability to retrieve and aggregate information; VTC-Reasoning, which requires models to infer latent associations to locate facts with minimal lexical overlap; and VTC-Memory, which measures comprehensive question answering within long-term dialogue memory. Furthermore, we establish the VTCBench-Wild to simulate diverse input scenarios.We comprehensively evaluate leading open-source and proprietary models on our benchmarks. The results indicate that, despite being able to decode textual information (e.g., OCR) well, most VLMs exhibit a surprisingly poor long-context understanding ability with VTC-processed information, failing to capture long associations or dependencies in the context.This study provides a deep understanding of VTC and serves as a foundation for designing more efficient and scalable VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15649
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
Zhao, Hongbo
Wang, Meng
Zhu, Fei
Liu, Wenzhuo
Ni, Bolin
Zeng, Fanhu
Meng, Gaofeng
Zhang, Zhaoxiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
The computational and memory overheads associated with expanding the context window of LLMs severely limit their scalability. A noteworthy solution is vision-text compression (VTC), exemplified by frameworks like DeepSeek-OCR and Glyph, which convert long texts into dense 2D visual representations, thereby achieving token compression ratios of 3x-20x. However, the impact of this high information density on the core long-context capabilities of vision-language models (VLMs) remains under-investigated. To address this gap, we introduce the first benchmark for VTC and systematically assess the performance of VLMs across three long-context understanding settings: VTC-Retrieval, which evaluates the model's ability to retrieve and aggregate information; VTC-Reasoning, which requires models to infer latent associations to locate facts with minimal lexical overlap; and VTC-Memory, which measures comprehensive question answering within long-term dialogue memory. Furthermore, we establish the VTCBench-Wild to simulate diverse input scenarios.We comprehensively evaluate leading open-source and proprietary models on our benchmarks. The results indicate that, despite being able to decode textual information (e.g., OCR) well, most VLMs exhibit a surprisingly poor long-context understanding ability with VTC-processed information, failing to capture long associations or dependencies in the context.This study provides a deep understanding of VTC and serves as a foundation for designing more efficient and scalable VLMs.
title VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.15649