MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sheikh, Tooba Tehreem, Lahoud, Jean, Anwer, Rao Muhammad, Khan, Fahad Shahbaz, Khan, Salman, Cholakkal, Hisham
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911286684549120
author Sheikh, Tooba Tehreem
Lahoud, Jean
Anwer, Rao Muhammad
Khan, Fahad Shahbaz
Khan, Salman
Cholakkal, Hisham
author_facet Sheikh, Tooba Tehreem
Lahoud, Jean
Anwer, Rao Muhammad
Khan, Fahad Shahbaz
Khan, Salman
Cholakkal, Hisham
contents Traditional object detection models in medical imaging operate within a closed-set paradigm, limiting their ability to detect objects of novel labels. Open-vocabulary object detection (OVOD) addresses this limitation but remains underexplored in medical imaging due to dataset scarcity and weak text-image alignment. To bridge this gap, we introduce MedROV, the first Real-time Open Vocabulary detection model for medical imaging. To enable open-vocabulary learning, we curate a large-scale dataset, Omnis, with 600K detection samples across nine imaging modalities and introduce a pseudo-labeling strategy to handle missing annotations from multi-source datasets. Additionally, we enhance generalization by incorporating knowledge from a large pre-trained foundation model. By leveraging contrastive learning and cross-modal representations, MedROV effectively detects both known and novel structures. Experimental results demonstrate that MedROV outperforms the previous state-of-the-art foundation model for medical image detection with an average absolute improvement of 40 mAP50, and surpasses closed-set detectors by more than 3 mAP50, while running at 70 FPS, setting a new benchmark in medical detection. Our source code, dataset, and trained model are available at https://github.com/toobatehreem/MedROV.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20650
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
Sheikh, Tooba Tehreem
Lahoud, Jean
Anwer, Rao Muhammad
Khan, Fahad Shahbaz
Khan, Salman
Cholakkal, Hisham
Computer Vision and Pattern Recognition
Artificial Intelligence
Traditional object detection models in medical imaging operate within a closed-set paradigm, limiting their ability to detect objects of novel labels. Open-vocabulary object detection (OVOD) addresses this limitation but remains underexplored in medical imaging due to dataset scarcity and weak text-image alignment. To bridge this gap, we introduce MedROV, the first Real-time Open Vocabulary detection model for medical imaging. To enable open-vocabulary learning, we curate a large-scale dataset, Omnis, with 600K detection samples across nine imaging modalities and introduce a pseudo-labeling strategy to handle missing annotations from multi-source datasets. Additionally, we enhance generalization by incorporating knowledge from a large pre-trained foundation model. By leveraging contrastive learning and cross-modal representations, MedROV effectively detects both known and novel structures. Experimental results demonstrate that MedROV outperforms the previous state-of-the-art foundation model for medical image detection with an average absolute improvement of 40 mAP50, and surpasses closed-set detectors by more than 3 mAP50, while running at 70 FPS, setting a new benchmark in medical detection. Our source code, dataset, and trained model are available at https://github.com/toobatehreem/MedROV.
title MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.20650