When LLaVA Meets Objects: Token Composition for Vision-Language-Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jahagirdar, Soumya, Bousselham, Walid, Kukleva, Anna, Kuehne, Hilde
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908822250979328
author Jahagirdar, Soumya
Bousselham, Walid
Kukleva, Anna
Kuehne, Hilde
author_facet Jahagirdar, Soumya
Bousselham, Walid
Kukleva, Anna
Kuehne, Hilde
contents Current autoregressive Vision Language Models (VLMs) usually rely on a large number of visual tokens to represent images, resulting in a need for more compute especially at inference time. To address this problem, we propose Mask-LLaVA, a framework that leverages different levels of visual features to create a compact yet information-rich visual representation for autoregressive VLMs. Namely, we combine mask-based object representations together with global tokens and local patch tokens. While all tokens are used during training, it shows that the resulting model can flexibly drop especially the number of mask-based object-tokens at test time, allowing to adapt the number of tokens during inference without the need to retrain the model and without a significant drop in performance. We evaluate the proposed approach on a suite of standard benchmarks showing results competitive to current token efficient methods and comparable to the original LLaVA baseline using only a fraction of visual tokens. Our analysis demonstrates that combining multi-level features enables efficient learning with fewer tokens while allowing dynamic token selection at test time for good performance.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04864
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When LLaVA Meets Objects: Token Composition for Vision-Language-Models
Jahagirdar, Soumya
Bousselham, Walid
Kukleva, Anna
Kuehne, Hilde
Computer Vision and Pattern Recognition
Current autoregressive Vision Language Models (VLMs) usually rely on a large number of visual tokens to represent images, resulting in a need for more compute especially at inference time. To address this problem, we propose Mask-LLaVA, a framework that leverages different levels of visual features to create a compact yet information-rich visual representation for autoregressive VLMs. Namely, we combine mask-based object representations together with global tokens and local patch tokens. While all tokens are used during training, it shows that the resulting model can flexibly drop especially the number of mask-based object-tokens at test time, allowing to adapt the number of tokens during inference without the need to retrain the model and without a significant drop in performance. We evaluate the proposed approach on a suite of standard benchmarks showing results competitive to current token efficient methods and comparable to the original LLaVA baseline using only a fraction of visual tokens. Our analysis demonstrates that combining multi-level features enables efficient learning with fewer tokens while allowing dynamic token selection at test time for good performance.
title When LLaVA Meets Objects: Token Composition for Vision-Language-Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.04864