How Much Information Can a Vision Token Hold? A Scaling Law for Recognition Limits in VLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhuang, Shuxin, Liang, Zi, Yu, Runsheng, Li, Hongzong, Feng, Rong, Tang, Shiqin, Zhang, Youzhi
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914303270977536
author Zhuang, Shuxin
Liang, Zi
Yu, Runsheng
Li, Hongzong
Feng, Rong
Tang, Shiqin
Zhang, Youzhi
author_facet Zhuang, Shuxin
Liang, Zi
Yu, Runsheng
Li, Hongzong
Feng, Rong
Tang, Shiqin
Zhang, Youzhi
contents Recent vision-centric approaches have made significant strides in long-context modeling. Represented by DeepSeek-OCR, these models encode rendered text into continuous vision tokens, achieving high compression rates without sacrificing recognition precision. However, viewing the vision encoder as a lossy channel with finite representational capacity raises a fundamental question: what is the information upper bound of visual tokens? To investigate this limit, we conduct controlled stress tests by progressively increasing the information quantity (character count) within an image. We observe a distinct phase-transition phenomenon characterized by three regimes: a near-perfect Stable Phase, an Instability Phase marked by increased error variance, and a total Collapse Phase. We analyze the mechanical origins of these transitions and identify key factors. Furthermore, we formulate a probabilistic scaling law that unifies average vision token load and visual density into a latent difficulty metric. Extensive experiments across various Vision-Language Models demonstrate the universality of this scaling law, providing critical empirical guidance for optimizing the efficiency-accuracy trade-off in visual context compression.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02539
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Much Information Can a Vision Token Hold? A Scaling Law for Recognition Limits in VLMs
Zhuang, Shuxin
Liang, Zi
Yu, Runsheng
Li, Hongzong
Feng, Rong
Tang, Shiqin
Zhang, Youzhi
Machine Learning
Computer Vision and Pattern Recognition
Recent vision-centric approaches have made significant strides in long-context modeling. Represented by DeepSeek-OCR, these models encode rendered text into continuous vision tokens, achieving high compression rates without sacrificing recognition precision. However, viewing the vision encoder as a lossy channel with finite representational capacity raises a fundamental question: what is the information upper bound of visual tokens? To investigate this limit, we conduct controlled stress tests by progressively increasing the information quantity (character count) within an image. We observe a distinct phase-transition phenomenon characterized by three regimes: a near-perfect Stable Phase, an Instability Phase marked by increased error variance, and a total Collapse Phase. We analyze the mechanical origins of these transitions and identify key factors. Furthermore, we formulate a probabilistic scaling law that unifies average vision token load and visual density into a latent difficulty metric. Extensive experiments across various Vision-Language Models demonstrate the universality of this scaling law, providing critical empirical guidance for optimizing the efficiency-accuracy trade-off in visual context compression.
title How Much Information Can a Vision Token Hold? A Scaling Law for Recognition Limits in VLMs
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.02539