How Much Information Can a Vision Token Hold? A Scaling Law for Recognition Limits in VLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhuang, Shuxin, Liang, Zi, Yu, Runsheng, Li, Hongzong, Feng, Rong, Tang, Shiqin, Zhang, Youzhi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Particle Dynamics for Latent-Variable Energy-Based Models
di: Tang, Shiqin, et al.
Pubblicazione: (2025)
di: Tang, Shiqin, et al.
Pubblicazione: (2025)
Adversary-Free Counterfactual Prediction via Information-Regularized Representations
di: Tang, Shiqin, et al.
Pubblicazione: (2025)
di: Tang, Shiqin, et al.
Pubblicazione: (2025)
How Vulnerable Are Edge LLMs?
di: Ding, Ao, et al.
Pubblicazione: (2026)
di: Ding, Ao, et al.
Pubblicazione: (2026)
Balancing Efficiency and Fairness: An Iterative Exchange Framework for Multi-UAV Cooperative Path Planning
di: Li, Hongzong, et al.
Pubblicazione: (2025)
di: Li, Hongzong, et al.
Pubblicazione: (2025)
Image Recognition with Vision and Language Embeddings of VLMs
di: Volkov, Illia, et al.
Pubblicazione: (2025)
di: Volkov, Illia, et al.
Pubblicazione: (2025)
Tree-Based Stochastic Optimization for Solving Large-Scale Urban Network Security Games
di: Zhuang, Shuxin, et al.
Pubblicazione: (2025)
di: Zhuang, Shuxin, et al.
Pubblicazione: (2025)
Towards Lossless Ultimate Vision Token Compression for VLMs
di: Zheng, Dehua, et al.
Pubblicazione: (2025)
di: Zheng, Dehua, et al.
Pubblicazione: (2025)
How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions
di: Bilmes, Jeff A., et al.
Pubblicazione: (2026)
di: Bilmes, Jeff A., et al.
Pubblicazione: (2026)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
di: Chen, Siyi, et al.
Pubblicazione: (2026)
di: Chen, Siyi, et al.
Pubblicazione: (2026)
Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs
di: Zhang, Song, et al.
Pubblicazione: (2026)
di: Zhang, Song, et al.
Pubblicazione: (2026)
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
di: Fang, Zhengyao, et al.
Pubblicazione: (2026)
di: Fang, Zhengyao, et al.
Pubblicazione: (2026)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
di: Wang, Feng, et al.
Pubblicazione: (2025)
di: Wang, Feng, et al.
Pubblicazione: (2025)
Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
di: Zhong, Yuan, et al.
Pubblicazione: (2025)
di: Zhong, Yuan, et al.
Pubblicazione: (2025)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
di: Tan, Zhangyun, et al.
Pubblicazione: (2026)
di: Tan, Zhangyun, et al.
Pubblicazione: (2026)
DeepSORT-Driven Visual Tracking Approach for Gesture Recognition in Interactive Systems
di: Zhang, Tong, et al.
Pubblicazione: (2025)
di: Zhang, Tong, et al.
Pubblicazione: (2025)
Guess the Unified Model: How Much Can We Recover from Generated Images?
di: Cekinmez, Jasin, et al.
Pubblicazione: (2026)
di: Cekinmez, Jasin, et al.
Pubblicazione: (2026)
ResWCAE: Biometric Pattern Image Denoising Using Residual Wavelet-Conditioned Autoencoder
di: Liang, Youzhi, et al.
Pubblicazione: (2023)
di: Liang, Youzhi, et al.
Pubblicazione: (2023)
Stateful Token Reduction for Long-Video Hybrid VLMs
di: Jiang, Jindong, et al.
Pubblicazione: (2026)
di: Jiang, Jindong, et al.
Pubblicazione: (2026)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
di: Gao, Yiling, et al.
Pubblicazione: (2026)
di: Gao, Yiling, et al.
Pubblicazione: (2026)
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
di: Xuan, Weihao, et al.
Pubblicazione: (2025)
di: Xuan, Weihao, et al.
Pubblicazione: (2025)
The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling
di: Shiba, Takuya
Pubblicazione: (2026)
di: Shiba, Takuya
Pubblicazione: (2026)
InnerGS: Internal Scenes Reconstruction and Segmentation via Factorized 3D Gaussian Splatting
di: Liang, Shuxin, et al.
Pubblicazione: (2025)
di: Liang, Shuxin, et al.
Pubblicazione: (2025)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
di: Yu, Shoubin, et al.
Pubblicazione: (2026)
di: Yu, Shoubin, et al.
Pubblicazione: (2026)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
di: Clark, Christopher, et al.
Pubblicazione: (2026)
di: Clark, Christopher, et al.
Pubblicazione: (2026)
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
di: Zhang, Gengyuan, et al.
Pubblicazione: (2023)
di: Zhang, Gengyuan, et al.
Pubblicazione: (2023)
Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration
di: Zeng, Fanhu, et al.
Pubblicazione: (2025)
di: Zeng, Fanhu, et al.
Pubblicazione: (2025)
Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding
di: Tang, Yutao, et al.
Pubblicazione: (2025)
di: Tang, Yutao, et al.
Pubblicazione: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
di: Lu, Meng, et al.
Pubblicazione: (2025)
di: Lu, Meng, et al.
Pubblicazione: (2025)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
ApET: Approximation-Error Guided Token Compression for Efficient VLMs
di: Ma, Qiankun, et al.
Pubblicazione: (2026)
di: Ma, Qiankun, et al.
Pubblicazione: (2026)
Enhancing Action Recognition from Low-Quality Skeleton Data via Part-Level Knowledge Distillation
di: Liu, Cuiwei, et al.
Pubblicazione: (2024)
di: Liu, Cuiwei, et al.
Pubblicazione: (2024)
TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
di: Zhao, Xuanle, et al.
Pubblicazione: (2025)
di: Zhao, Xuanle, et al.
Pubblicazione: (2025)
Spectral Vision Transformer for Efficient Tokenization with Limited Data
di: Roberts, Alexandra G., et al.
Pubblicazione: (2026)
di: Roberts, Alexandra G., et al.
Pubblicazione: (2026)
T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition
di: Yeh, Chen, et al.
Pubblicazione: (2024)
di: Yeh, Chen, et al.
Pubblicazione: (2024)
How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion Models
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
TopoBind: Multi-Modal Prediction of Antibody-Antigen Binding Free Energy via Sequence Embeddings and Structural Topology
di: Yu, Ciyuan, et al.
Pubblicazione: (2025)
di: Yu, Ciyuan, et al.
Pubblicazione: (2025)
ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification
di: He, Yefei, et al.
Pubblicazione: (2024)
di: He, Yefei, et al.
Pubblicazione: (2024)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
di: Zhang, Jianrui, et al.
Pubblicazione: (2026)
di: Zhang, Jianrui, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Particle Dynamics for Latent-Variable Energy-Based Models
di: Tang, Shiqin, et al.
Pubblicazione: (2025) -
Adversary-Free Counterfactual Prediction via Information-Regularized Representations
di: Tang, Shiqin, et al.
Pubblicazione: (2025) -
How Vulnerable Are Edge LLMs?
di: Ding, Ao, et al.
Pubblicazione: (2026) -
Balancing Efficiency and Fairness: An Iterative Exchange Framework for Multi-UAV Cooperative Path Planning
di: Li, Hongzong, et al.
Pubblicazione: (2025) -
Image Recognition with Vision and Language Embeddings of VLMs
di: Volkov, Illia, et al.
Pubblicazione: (2025)