OSR-ViT: A Simple and Modular Framework for Open-Set Object Detection and Discovery

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Inkawhich, Matthew, Inkawhich, Nathan, Yang, Hao, Zhang, Jingyang, Linderman, Randolph, Chen, Yiran
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929315894001664
author Inkawhich, Matthew
Inkawhich, Nathan
Yang, Hao
Zhang, Jingyang
Linderman, Randolph
Chen, Yiran
author_facet Inkawhich, Matthew
Inkawhich, Nathan
Yang, Hao
Zhang, Jingyang
Linderman, Randolph
Chen, Yiran
contents An object detector's ability to detect and flag \textit{novel} objects during open-world deployments is critical for many real-world applications. Unfortunately, much of the work in open object detection today is disjointed and fails to adequately address applications that prioritize unknown object recall \textit{in addition to} known-class accuracy. To close this gap, we present a new task called Open-Set Object Detection and Discovery (OSODD) and as a solution propose the Open-Set Regions with ViT features (OSR-ViT) detection framework. OSR-ViT combines a class-agnostic proposal network with a powerful ViT-based classifier. Its modular design simplifies optimization and allows users to easily swap proposal solutions and feature extractors to best suit their application. Using our multifaceted evaluation protocol, we show that OSR-ViT obtains performance levels that far exceed state-of-the-art supervised methods. Our method also excels in low-data settings, outperforming supervised baselines using a fraction of the training data.
format Preprint
id arxiv_https___arxiv_org_abs_2404_10865
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OSR-ViT: A Simple and Modular Framework for Open-Set Object Detection and Discovery
Inkawhich, Matthew
Inkawhich, Nathan
Yang, Hao
Zhang, Jingyang
Linderman, Randolph
Chen, Yiran
Computer Vision and Pattern Recognition
An object detector's ability to detect and flag \textit{novel} objects during open-world deployments is critical for many real-world applications. Unfortunately, much of the work in open object detection today is disjointed and fails to adequately address applications that prioritize unknown object recall \textit{in addition to} known-class accuracy. To close this gap, we present a new task called Open-Set Object Detection and Discovery (OSODD) and as a solution propose the Open-Set Regions with ViT features (OSR-ViT) detection framework. OSR-ViT combines a class-agnostic proposal network with a powerful ViT-based classifier. Its modular design simplifies optimization and allows users to easily swap proposal solutions and feature extractors to best suit their application. Using our multifaceted evaluation protocol, we show that OSR-ViT obtains performance levels that far exceed state-of-the-art supervised methods. Our method also excels in low-data settings, outperforming supervised baselines using a fraction of the training data.
title OSR-ViT: A Simple and Modular Framework for Open-Set Object Detection and Discovery
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.10865