DSAA: Dual-Stage Attribute Activation for Fine-grained Open Vocabulary Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Donghong, Lin, Endian, Liu, Hanqing, Liu, Mingjie, Cui, Luoping, Yang, Zhao, Zhu, Chuang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914617394987008
author Jiang, Donghong
Lin, Endian
Liu, Hanqing
Liu, Mingjie
Cui, Luoping
Yang, Zhao
Zhu, Chuang
author_facet Jiang, Donghong
Lin, Endian
Liu, Hanqing
Liu, Mingjie
Cui, Luoping
Yang, Zhao
Zhu, Chuang
contents Open-Vocabulary Object Detection (OVD) models break the limitations of closed-set detection, enabling the identification of unseen categories through natural language prompts. However, they exhibit notable limitations in fine-grained detection tasks involving attributes like color, material, and texture. We attribute this performance bottleneck in OVD models to a core issue: when category signals dominate, OVD models tend to marginalize attribute information during inference. This leads to incorrect binding between attributes and target objects. To address this, we propose the Dual-Stage Attribute Activation (DSAA) framework, which enhances fine-grained detection capabilities by strengthening attribute semantics at two critical stages. In the text embedding stage, we employ Attribute Prefix Adapter (APA) module to generate attribute prefixes that inject explicit attribute priors. To further amplify the influence of these attributes, our Key/Value (K/V) Modulator module then intervenes during the BERT encoding phase, selectively enhancing the Key and Value vectors of the corresponding attribute tokens. In addition, we introduce an attribute-aware contrastive loss to improve discrimination among same-category instances with different attributes during training. Experimental results on the FG-OVD benchmark demonstrate the effectiveness of our method across various mainstream open-vocabulary models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18023
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DSAA: Dual-Stage Attribute Activation for Fine-grained Open Vocabulary Detection
Jiang, Donghong
Lin, Endian
Liu, Hanqing
Liu, Mingjie
Cui, Luoping
Yang, Zhao
Zhu, Chuang
Computer Vision and Pattern Recognition
Open-Vocabulary Object Detection (OVD) models break the limitations of closed-set detection, enabling the identification of unseen categories through natural language prompts. However, they exhibit notable limitations in fine-grained detection tasks involving attributes like color, material, and texture. We attribute this performance bottleneck in OVD models to a core issue: when category signals dominate, OVD models tend to marginalize attribute information during inference. This leads to incorrect binding between attributes and target objects. To address this, we propose the Dual-Stage Attribute Activation (DSAA) framework, which enhances fine-grained detection capabilities by strengthening attribute semantics at two critical stages. In the text embedding stage, we employ Attribute Prefix Adapter (APA) module to generate attribute prefixes that inject explicit attribute priors. To further amplify the influence of these attributes, our Key/Value (K/V) Modulator module then intervenes during the BERT encoding phase, selectively enhancing the Key and Value vectors of the corresponding attribute tokens. In addition, we introduce an attribute-aware contrastive loss to improve discrimination among same-category instances with different attributes during training. Experimental results on the FG-OVD benchmark demonstrate the effectiveness of our method across various mainstream open-vocabulary models.
title DSAA: Dual-Stage Attribute Activation for Fine-grained Open Vocabulary Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.18023