Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Robicheaux, Peter, Popov, Matvei, Madan, Anish, Robinson, Isaac, Nelson, Joseph, Ramanan, Deva, Peri, Neehar
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914108553560064
author Robicheaux, Peter
Popov, Matvei
Madan, Anish
Robinson, Isaac
Nelson, Joseph
Ramanan, Deva
Peri, Neehar
author_facet Robicheaux, Peter
Popov, Matvei
Madan, Anish
Robinson, Isaac
Nelson, Joseph
Ramanan, Deva
Peri, Neehar
contents Vision-language models (VLMs) trained on internet-scale data achieve remarkable zero-shot detection performance on common objects like car, truck, and pedestrian. However, state-of-the-art models still struggle to generalize to out-of-distribution classes, tasks and imaging modalities not typically found in their pre-training. Rather than simply re-training VLMs on more visual data, we argue that one should align VLMs to new concepts with annotation instructions containing a few visual examples and rich textual descriptions. To this end, we introduce Roboflow100-VL, a large-scale collection of 100 multi-modal object detection datasets with diverse concepts not commonly found in VLM pre-training. We evaluate state-of-the-art models on our benchmark in zero-shot, few-shot, semi-supervised, and fully-supervised settings, allowing for comparison across data regimes. Notably, we find that VLMs like GroundingDINO and Qwen2.5-VL achieve less than 2% zero-shot accuracy on challenging medical imaging datasets within Roboflow100-VL, demonstrating the need for few-shot concept alignment. Lastly, we discuss our recent CVPR 2025 Foundational FSOD competition and share insights from the community. Notably, the winning team significantly outperforms our baseline by 17 mAP! Our code and dataset are available at https://github.com/roboflow/rf100-vl and https://universe.roboflow.com/rf100-vl/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20612
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
Robicheaux, Peter
Popov, Matvei
Madan, Anish
Robinson, Isaac
Nelson, Joseph
Ramanan, Deva
Peri, Neehar
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Vision-language models (VLMs) trained on internet-scale data achieve remarkable zero-shot detection performance on common objects like car, truck, and pedestrian. However, state-of-the-art models still struggle to generalize to out-of-distribution classes, tasks and imaging modalities not typically found in their pre-training. Rather than simply re-training VLMs on more visual data, we argue that one should align VLMs to new concepts with annotation instructions containing a few visual examples and rich textual descriptions. To this end, we introduce Roboflow100-VL, a large-scale collection of 100 multi-modal object detection datasets with diverse concepts not commonly found in VLM pre-training. We evaluate state-of-the-art models on our benchmark in zero-shot, few-shot, semi-supervised, and fully-supervised settings, allowing for comparison across data regimes. Notably, we find that VLMs like GroundingDINO and Qwen2.5-VL achieve less than 2% zero-shot accuracy on challenging medical imaging datasets within Roboflow100-VL, demonstrating the need for few-shot concept alignment. Lastly, we discuss our recent CVPR 2025 Foundational FSOD competition and share insights from the community. Notably, the winning team significantly outperforms our baseline by 17 mAP! Our code and dataset are available at https://github.com/roboflow/rf100-vl and https://universe.roboflow.com/rf100-vl/.
title Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.20612