ART: Adaptive Relation Tuning for Generalized Relation Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sudhakaran, Gopika, Shindo, Hikaru, Schramowski, Patrick, Schaub-Meyer, Simone, Kersting, Kristian, Roth, Stefan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918118230589440
author Sudhakaran, Gopika
Shindo, Hikaru
Schramowski, Patrick
Schaub-Meyer, Simone
Kersting, Kristian
Roth, Stefan
author_facet Sudhakaran, Gopika
Shindo, Hikaru
Schramowski, Patrick
Schaub-Meyer, Simone
Kersting, Kristian
Roth, Stefan
contents Visual relation detection (VRD) is the task of identifying the relationships between objects in a scene. VRD models trained solely on relation detection data struggle to generalize beyond the relations on which they are trained. While prompt tuning has been used to adapt vision-language models (VLMs) for VRD, it uses handcrafted prompts and struggles with novel or complex relations. We argue that instruction tuning offers a more effective solution by fine-tuning VLMs on diverse instructional data. We thus introduce ART, an Adaptive Relation Tuning framework that adapts VLMs for VRD through instruction tuning and strategic instance selection. By converting VRD datasets into an instruction tuning format and employing an adaptive sampling algorithm, ART directs the VLM to focus on informative relations while maintaining generalizability. Specifically, we focus on the relation classification, where subject-object boxes are given and the model predicts the predicate between them. We tune on a held-in set and evaluate across multiple held-out datasets of varying complexity. Our approach strongly improves over its baselines and can infer unseen relation concepts, a capability absent in mainstream VRD methods. We demonstrate ART's practical value by using the predicted relations for segmenting complex scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23543
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ART: Adaptive Relation Tuning for Generalized Relation Prediction
Sudhakaran, Gopika
Shindo, Hikaru
Schramowski, Patrick
Schaub-Meyer, Simone
Kersting, Kristian
Roth, Stefan
Computer Vision and Pattern Recognition
Artificial Intelligence
Visual relation detection (VRD) is the task of identifying the relationships between objects in a scene. VRD models trained solely on relation detection data struggle to generalize beyond the relations on which they are trained. While prompt tuning has been used to adapt vision-language models (VLMs) for VRD, it uses handcrafted prompts and struggles with novel or complex relations. We argue that instruction tuning offers a more effective solution by fine-tuning VLMs on diverse instructional data. We thus introduce ART, an Adaptive Relation Tuning framework that adapts VLMs for VRD through instruction tuning and strategic instance selection. By converting VRD datasets into an instruction tuning format and employing an adaptive sampling algorithm, ART directs the VLM to focus on informative relations while maintaining generalizability. Specifically, we focus on the relation classification, where subject-object boxes are given and the model predicts the predicate between them. We tune on a held-in set and evaluate across multiple held-out datasets of varying complexity. Our approach strongly improves over its baselines and can infer unseen relation concepts, a capability absent in mainstream VRD methods. We demonstrate ART's practical value by using the predicted relations for segmenting complex scenes.
title ART: Adaptive Relation Tuning for Generalized Relation Prediction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.23543