From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Rongjie, Zhang, Songyang, Lin, Dahua, Chen, Kai, He, Xuming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910420703379456
author Li, Rongjie
Zhang, Songyang
Lin, Dahua
Chen, Kai
He, Xuming
author_facet Li, Rongjie
Zhang, Songyang
Lin, Dahua
Chen, Kai
He, Xuming
contents Scene graph generation (SGG) aims to parse a visual scene into an intermediate graph representation for downstream reasoning tasks. Despite recent advancements, existing methods struggle to generate scene graphs with novel visual relation concepts. To address this challenge, we introduce a new open-vocabulary SGG framework based on sequence generation. Our framework leverages vision-language pre-trained models (VLM) by incorporating an image-to-graph generation paradigm. Specifically, we generate scene graph sequences via image-to-text generation with VLM and then construct scene graphs from these sequences. By doing so, we harness the strong capabilities of VLM for open-vocabulary SGG and seamlessly integrate explicit relational modeling for enhancing the VL tasks. Experimental results demonstrate that our design not only achieves superior performance with an open vocabulary but also enhances downstream vision-language task performance through explicit relation modeling knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2404_00906
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
Li, Rongjie
Zhang, Songyang
Lin, Dahua
Chen, Kai
He, Xuming
Computer Vision and Pattern Recognition
Scene graph generation (SGG) aims to parse a visual scene into an intermediate graph representation for downstream reasoning tasks. Despite recent advancements, existing methods struggle to generate scene graphs with novel visual relation concepts. To address this challenge, we introduce a new open-vocabulary SGG framework based on sequence generation. Our framework leverages vision-language pre-trained models (VLM) by incorporating an image-to-graph generation paradigm. Specifically, we generate scene graph sequences via image-to-text generation with VLM and then construct scene graphs from these sequences. By doing so, we harness the strong capabilities of VLM for open-vocabulary SGG and seamlessly integrate explicit relational modeling for enhancing the VL tasks. Experimental results demonstrate that our design not only achieves superior performance with an open vocabulary but also enhances downstream vision-language task performance through explicit relation modeling knowledge.
title From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.00906