Self-Improving Small Object Grounding in LVLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Tianze, Shi, Yucheng, Sun, Ruitong, Liu, Ninghao, Sun, Jin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918534513164288
author Yang, Tianze
Shi, Yucheng
Sun, Ruitong
Liu, Ninghao
Sun, Jin
author_facet Yang, Tianze
Shi, Yucheng
Sun, Ruitong
Liu, Ninghao
Sun, Jin
contents Can internal attention patterns in Large Vision Language Models (LVLMs) identify reliable small-object boxes without fine-tuning? In this work, we provide an affirmative answer. Attention structure in LVLMs encodes grounding quality-a lightweight IoU regressor trained solely on attention maps achieves strong IoU prediction (Pearson r > 0.67). This regressor powers the regressor-based variant of our Attention-based Candidate Selection (ACS) framework, called ACS-Learned, which selects the best box from multiple sampled candidates to improve object grounding. By analyzing what the regressor learns, we reveal which transformer layers and heads are most critical and derive ACS-Free: a training-free selector that ranks candidates by attention entropy on these discriminative heads, with no learned component at inference. Experiments on COCO and Objects365 demonstrate up to 19% self-improvement on small object localization, with ACS-Free ranking best among all training-free methods, demonstrating that useful attention structure improves both localization reliability and interpretability in LVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01612
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Self-Improving Small Object Grounding in LVLMs
Yang, Tianze
Shi, Yucheng
Sun, Ruitong
Liu, Ninghao
Sun, Jin
Computer Vision and Pattern Recognition
Machine Learning
Can internal attention patterns in Large Vision Language Models (LVLMs) identify reliable small-object boxes without fine-tuning? In this work, we provide an affirmative answer. Attention structure in LVLMs encodes grounding quality-a lightweight IoU regressor trained solely on attention maps achieves strong IoU prediction (Pearson r > 0.67). This regressor powers the regressor-based variant of our Attention-based Candidate Selection (ACS) framework, called ACS-Learned, which selects the best box from multiple sampled candidates to improve object grounding. By analyzing what the regressor learns, we reveal which transformer layers and heads are most critical and derive ACS-Free: a training-free selector that ranks candidates by attention entropy on these discriminative heads, with no learned component at inference. Experiments on COCO and Objects365 demonstrate up to 19% self-improvement on small object localization, with ACS-Free ranking best among all training-free methods, demonstrating that useful attention structure improves both localization reliability and interpretability in LVLMs.
title Self-Improving Small Object Grounding in LVLMs
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2606.01612