Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Vellamcheti, Shanmukha, Kundu, Sanjoy, Aakur, Sathyanarayanan N.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913881059753984
author Vellamcheti, Shanmukha
Kundu, Sanjoy
Aakur, Sathyanarayanan N.
author_facet Vellamcheti, Shanmukha
Kundu, Sanjoy
Aakur, Sathyanarayanan N.
contents Understanding relationships between objects is central to visual intelligence, with applications in embodied AI, assistive systems, and scene understanding. Yet, most visual relationship detection (VRD) models rely on a fixed predicate set, limiting their generalization to novel interactions. A key challenge is the inability to visually ground semantically plausible, but unannotated, relationships hypothesized from external knowledge. This work introduces an iterative visual grounding framework that leverages large language models (LLMs) as structured relational priors. Inspired by expectation-maximization (EM), our method alternates between generating candidate scene graphs from detected objects using an LLM (expectation) and training a visual model to align these hypotheses with perceptual evidence (maximization). This process bootstraps relational understanding beyond annotated data and enables generalization to unseen predicates. Additionally, we introduce a new benchmark for open-world VRD on Visual Genome with 21 held-out predicates and evaluate under three settings: seen, unseen, and mixed. Our model outperforms LLM-only, few-shot, and debiased baselines, achieving mean recall (mR@50) of 15.9, 13.1, and 11.7 on predicate classification on these three sets. These results highlight the promise of grounded LLM priors for scalable open-world visual understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05651
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
Vellamcheti, Shanmukha
Kundu, Sanjoy
Aakur, Sathyanarayanan N.
Computer Vision and Pattern Recognition
Understanding relationships between objects is central to visual intelligence, with applications in embodied AI, assistive systems, and scene understanding. Yet, most visual relationship detection (VRD) models rely on a fixed predicate set, limiting their generalization to novel interactions. A key challenge is the inability to visually ground semantically plausible, but unannotated, relationships hypothesized from external knowledge. This work introduces an iterative visual grounding framework that leverages large language models (LLMs) as structured relational priors. Inspired by expectation-maximization (EM), our method alternates between generating candidate scene graphs from detected objects using an LLM (expectation) and training a visual model to align these hypotheses with perceptual evidence (maximization). This process bootstraps relational understanding beyond annotated data and enables generalization to unseen predicates. Additionally, we introduce a new benchmark for open-world VRD on Visual Genome with 21 held-out predicates and evaluate under three settings: seen, unseen, and mixed. Our model outperforms LLM-only, few-shot, and debiased baselines, achieving mean recall (mR@50) of 15.9, 13.1, and 11.7 on predicate classification on these three sets. These results highlight the promise of grounded LLM priors for scalable open-world visual understanding.
title Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.05651