Beyond Referring Expressions: Scenario Comprehension Visual Grounding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: He, Ruozhen, Shah, Nisarg A., Dong, Qihua, Xiao, Zilin, Koo, Jaywon, Ordonez, Vicente
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917380985192448
author He, Ruozhen
Shah, Nisarg A.
Dong, Qihua
Xiao, Zilin
Koo, Jaywon
Ordonez, Vicente
author_facet He, Ruozhen
Shah, Nisarg A.
Dong, Qihua
Xiao, Zilin
Koo, Jaywon
Ordonez, Vicente
contents Existing visual grounding benchmarks primarily evaluate alignment between image regions and literal referring expressions, where models can often succeed by matching a prominent named category. We explore a complementary and more challenging setting of scenario-based visual grounding, where the target must be inferred from roles, intentions, and relational context rather than explicit naming. We introduce Referring Scenario Comprehension (RSC), a benchmark designed for this setting. The queries in this benchmark are paragraph-length texts that describe object roles, user goals, and contextual cues, including deliberate references to distractor objects that often require deep understanding to resolve. Each instance is annotated with interpretable difficulty tags for uniqueness, clutter, size, overlap, and position which expose distinct failure modes and support fine-grained analysis. RSC contains approximately 31k training examples, 4k in-domain test examples, and a 3k out-of-distribution split with unseen object categories. We further propose ScenGround, a curriculum reasoning method serving as a reference point for this setting, combining supervised warm-starting with difficulty-aware reinforcement learning. Experiments show that scenario-based queries expose systematic failures in current models that standard benchmarks do not reveal, and that curriculum training improves performance on challenging slices and transfers to standard benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02323
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Referring Expressions: Scenario Comprehension Visual Grounding
He, Ruozhen
Shah, Nisarg A.
Dong, Qihua
Xiao, Zilin
Koo, Jaywon
Ordonez, Vicente
Computer Vision and Pattern Recognition
Existing visual grounding benchmarks primarily evaluate alignment between image regions and literal referring expressions, where models can often succeed by matching a prominent named category. We explore a complementary and more challenging setting of scenario-based visual grounding, where the target must be inferred from roles, intentions, and relational context rather than explicit naming. We introduce Referring Scenario Comprehension (RSC), a benchmark designed for this setting. The queries in this benchmark are paragraph-length texts that describe object roles, user goals, and contextual cues, including deliberate references to distractor objects that often require deep understanding to resolve. Each instance is annotated with interpretable difficulty tags for uniqueness, clutter, size, overlap, and position which expose distinct failure modes and support fine-grained analysis. RSC contains approximately 31k training examples, 4k in-domain test examples, and a 3k out-of-distribution split with unseen object categories. We further propose ScenGround, a curriculum reasoning method serving as a reference point for this setting, combining supervised warm-starting with difficulty-aware reinforcement learning. Experiments show that scenario-based queries expose systematic failures in current models that standard benchmarks do not reveal, and that curriculum training improves performance on challenging slices and transfers to standard benchmarks.
title Beyond Referring Expressions: Scenario Comprehension Visual Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.02323