UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liao, Xinyao, Wei, Wei, Chen, Dangyang, Fu, Yuanyuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913643542609920
author Liao, Xinyao
Wei, Wei
Chen, Dangyang
Fu, Yuanyuan
author_facet Liao, Xinyao
Wei, Wei
Chen, Dangyang
Fu, Yuanyuan
contents Scene Graph Generation(SGG) is a scene understanding task that aims at identifying object entities and reasoning their relationships within a given image. In contrast to prevailing two-stage methods based on a large object detector (e.g., Faster R-CNN), one-stage methods integrate a fixed-size set of learnable queries to jointly reason relational triplets <subject, predicate, object>. This paradigm demonstrates robust performance with significantly reduced parameters and computational overhead. However, the challenge in one-stage methods stems from the issue of weak entanglement, wherein entities involved in relationships require both coupled features shared within triplets and decoupled visual features. Previous methods either adopt a single decoder for coupled triplet feature modeling or multiple decoders for separate visual feature extraction but fail to consider both. In this paper, we introduce UniQ, a Unified decoder with task-specific Queries architecture, where task-specific queries generate decoupled visual features for subjects, objects, and predicates respectively, and unified decoder enables coupled feature modeling within relational triplets. Experimental results on the Visual Genome dataset demonstrate that UniQ has superior performance to both one-stage and two-stage methods.
format Preprint
id arxiv_https___arxiv_org_abs_2501_05687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph Generation
Liao, Xinyao
Wei, Wei
Chen, Dangyang
Fu, Yuanyuan
Computer Vision and Pattern Recognition
I.2.10
Scene Graph Generation(SGG) is a scene understanding task that aims at identifying object entities and reasoning their relationships within a given image. In contrast to prevailing two-stage methods based on a large object detector (e.g., Faster R-CNN), one-stage methods integrate a fixed-size set of learnable queries to jointly reason relational triplets <subject, predicate, object>. This paradigm demonstrates robust performance with significantly reduced parameters and computational overhead. However, the challenge in one-stage methods stems from the issue of weak entanglement, wherein entities involved in relationships require both coupled features shared within triplets and decoupled visual features. Previous methods either adopt a single decoder for coupled triplet feature modeling or multiple decoders for separate visual feature extraction but fail to consider both. In this paper, we introduce UniQ, a Unified decoder with task-specific Queries architecture, where task-specific queries generate decoupled visual features for subjects, objects, and predicates respectively, and unified decoder enables coupled feature modeling within relational triplets. Experimental results on the Visual Genome dataset demonstrate that UniQ has superior performance to both one-stage and two-stage methods.
title UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph Generation
topic Computer Vision and Pattern Recognition
I.2.10
url https://arxiv.org/abs/2501.05687