Generating Enhanced Negatives for Training Language-Based Object Detectors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Shiyu, Zhao, Long, G, Vijay Kumar B., Suh, Yumin, Metaxas, Dimitris N., Chandraker, Manmohan, Schulter, Samuel
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910409161703424
author Zhao, Shiyu
Zhao, Long
G, Vijay Kumar B.
Suh, Yumin
Metaxas, Dimitris N.
Chandraker, Manmohan
Schulter, Samuel
author_facet Zhao, Shiyu
Zhao, Long
G, Vijay Kumar B.
Suh, Yumin
Metaxas, Dimitris N.
Chandraker, Manmohan
Schulter, Samuel
contents The recent progress in language-based open-vocabulary object detection can be largely attributed to finding better ways of leveraging large-scale data with free-form text annotations. Training such models with a discriminative objective function has proven successful, but requires good positive and negative samples. However, the free-form nature and the open vocabulary of object descriptions make the space of negatives extremely large. Prior works randomly sample negatives or use rule-based techniques to build them. In contrast, we propose to leverage the vast knowledge built into modern generative models to automatically build negatives that are more relevant to the original data. Specifically, we use large-language-models to generate negative text descriptions, and text-to-image diffusion models to also generate corresponding negative images. Our experimental analysis confirms the relevance of the generated negative data, and its use in language-based detectors improves performance on two complex benchmarks. Code is available at \url{https://github.com/xiaofeng94/Gen-Enhanced-Negs}.
format Preprint
id arxiv_https___arxiv_org_abs_2401_00094
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Generating Enhanced Negatives for Training Language-Based Object Detectors
Zhao, Shiyu
Zhao, Long
G, Vijay Kumar B.
Suh, Yumin
Metaxas, Dimitris N.
Chandraker, Manmohan
Schulter, Samuel
Computer Vision and Pattern Recognition
The recent progress in language-based open-vocabulary object detection can be largely attributed to finding better ways of leveraging large-scale data with free-form text annotations. Training such models with a discriminative objective function has proven successful, but requires good positive and negative samples. However, the free-form nature and the open vocabulary of object descriptions make the space of negatives extremely large. Prior works randomly sample negatives or use rule-based techniques to build them. In contrast, we propose to leverage the vast knowledge built into modern generative models to automatically build negatives that are more relevant to the original data. Specifically, we use large-language-models to generate negative text descriptions, and text-to-image diffusion models to also generate corresponding negative images. Our experimental analysis confirms the relevance of the generated negative data, and its use in language-based detectors improves performance on two complex benchmarks. Code is available at \url{https://github.com/xiaofeng94/Gen-Enhanced-Negs}.
title Generating Enhanced Negatives for Training Language-Based Object Detectors
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.00094