OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yu, Su, Xiangbo, Chen, Qiang, Zhang, Xinyu, Xi, Teng, Yao, Kun, Ding, Errui, Zhang, Gang, Wang, Jingdong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929421279035392
author Wang, Yu
Su, Xiangbo
Chen, Qiang
Zhang, Xinyu
Xi, Teng
Yao, Kun
Ding, Errui
Zhang, Gang
Wang, Jingdong
author_facet Wang, Yu
Su, Xiangbo
Chen, Qiang
Zhang, Xinyu
Xi, Teng
Yao, Kun
Ding, Errui
Zhang, Gang
Wang, Jingdong
contents Open-vocabulary object detection focusing on detecting novel categories guided by natural language. In this report, we propose Open-Vocabulary Light-Weighted Detection Transformer (OVLW-DETR), a deployment friendly open-vocabulary detector with strong performance and low latency. Building upon OVLW-DETR, we provide an end-to-end training recipe that transferring knowledge from vision-language model (VLM) to object detector with simple alignment. We align detector with the text encoder from VLM by replacing the fixed classification layer weights in detector with the class-name embeddings extracted from the text encoder. Without additional fusing module, OVLW-DETR is flexible and deployment friendly, making it easier to implement and modulate. improving the efficiency of interleaved attention computation. Experimental results demonstrate that the proposed approach is superior over existing real-time open-vocabulary detectors on standard Zero-Shot LVIS benchmark. Source code and pre-trained models are available at [https://github.com/Atten4Vis/LW-DETR].
format Preprint
id arxiv_https___arxiv_org_abs_2407_10655
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer
Wang, Yu
Su, Xiangbo
Chen, Qiang
Zhang, Xinyu
Xi, Teng
Yao, Kun
Ding, Errui
Zhang, Gang
Wang, Jingdong
Computer Vision and Pattern Recognition
Open-vocabulary object detection focusing on detecting novel categories guided by natural language. In this report, we propose Open-Vocabulary Light-Weighted Detection Transformer (OVLW-DETR), a deployment friendly open-vocabulary detector with strong performance and low latency. Building upon OVLW-DETR, we provide an end-to-end training recipe that transferring knowledge from vision-language model (VLM) to object detector with simple alignment. We align detector with the text encoder from VLM by replacing the fixed classification layer weights in detector with the class-name embeddings extracted from the text encoder. Without additional fusing module, OVLW-DETR is flexible and deployment friendly, making it easier to implement and modulate. improving the efficiency of interleaved attention computation. Experimental results demonstrate that the proposed approach is superior over existing real-time open-vocabulary detectors on standard Zero-Shot LVIS benchmark. Source code and pre-trained models are available at [https://github.com/Atten4Vis/LW-DETR].
title OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.10655