Enhancing Open-Vocabulary Object Detection through Multi-Level Fine-Grained Visual-Language Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tianyi, Simoulin, Antoine, Li, Kai, Lakdawala, Sana, Yu, Shiqing, Mittal, Arpit, Fu, Hongyu, Lin, Yu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912866278309888
author Zhang, Tianyi
Simoulin, Antoine
Li, Kai
Lakdawala, Sana
Yu, Shiqing
Mittal, Arpit
Fu, Hongyu
Lin, Yu
author_facet Zhang, Tianyi
Simoulin, Antoine
Li, Kai
Lakdawala, Sana
Yu, Shiqing
Mittal, Arpit
Fu, Hongyu
Lin, Yu
contents Traditional object detection systems are typically constrained to predefined categories, limiting their applicability in dynamic environments. In contrast, open-vocabulary object detection (OVD) enables the identification of objects from novel classes not present in the training set. Recent advances in visual-language modeling have led to significant progress of OVD. However, prior works face challenges in either adapting the single-scale image backbone from CLIP to the detection framework or ensuring robust visual-language alignment. We propose Visual-Language Detection (VLDet), a novel framework that revamps feature pyramid for fine-grained visual-language alignment, leading to improved OVD performance. With the VL-PUB module, VLDet effectively exploits the visual-language knowledge from CLIP and adapts the backbone for object detection through feature pyramid. In addition, we introduce the SigRPN block, which incorporates a sigmoid-based anchor-text contrastive alignment loss to improve detection of novel categories. Through extensive experiments, our approach achieves 58.7 AP for novel classes on COCO2017 and 24.8 AP on LVIS, surpassing all state-of-the-art methods and achieving significant improvements of 27.6% and 6.9%, respectively. Furthermore, VLDet also demonstrates superior zero-shot performance on closed-set object detection.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00531
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Enhancing Open-Vocabulary Object Detection through Multi-Level Fine-Grained Visual-Language Alignment
Zhang, Tianyi
Simoulin, Antoine
Li, Kai
Lakdawala, Sana
Yu, Shiqing
Mittal, Arpit
Fu, Hongyu
Lin, Yu
Computer Vision and Pattern Recognition
Traditional object detection systems are typically constrained to predefined categories, limiting their applicability in dynamic environments. In contrast, open-vocabulary object detection (OVD) enables the identification of objects from novel classes not present in the training set. Recent advances in visual-language modeling have led to significant progress of OVD. However, prior works face challenges in either adapting the single-scale image backbone from CLIP to the detection framework or ensuring robust visual-language alignment. We propose Visual-Language Detection (VLDet), a novel framework that revamps feature pyramid for fine-grained visual-language alignment, leading to improved OVD performance. With the VL-PUB module, VLDet effectively exploits the visual-language knowledge from CLIP and adapts the backbone for object detection through feature pyramid. In addition, we introduce the SigRPN block, which incorporates a sigmoid-based anchor-text contrastive alignment loss to improve detection of novel categories. Through extensive experiments, our approach achieves 58.7 AP for novel classes on COCO2017 and 24.8 AP on LVIS, surpassing all state-of-the-art methods and achieving significant improvements of 27.6% and 6.9%, respectively. Furthermore, VLDet also demonstrates superior zero-shot performance on closed-set object detection.
title Enhancing Open-Vocabulary Object Detection through Multi-Level Fine-Grained Visual-Language Alignment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.00531