MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ikezogwo, Wisdom O., Zhang, Kevin, Seyfioglu, Mehmet Saygin, Ghezloo, Fatemeh, Shapiro, Linda, Krishna, Ranjay
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908762715979776
author Ikezogwo, Wisdom O.
Zhang, Kevin
Seyfioglu, Mehmet Saygin
Ghezloo, Fatemeh
Shapiro, Linda
Krishna, Ranjay
author_facet Ikezogwo, Wisdom O.
Zhang, Kevin
Seyfioglu, Mehmet Saygin
Ghezloo, Fatemeh
Shapiro, Linda
Krishna, Ranjay
contents Multi-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury. To enable representation learning for medical images at scale, we turn to YouTube, a platform with a large reservoir of open-source medical pedagogical videos. We curate MedicalNarratives, a dataset 4.7M medical image-text pairs, with 1M samples containing dense annotations in the form of spatial traces (and bounding boxes), and 118K videos centered on the trace event (with aligned text), enabling spatiotemporal grounding beyond single frames. Similar to $\textit{think-aloud}$ studies where instructors speak while hovering their mouse cursor movements over relevant image regions, 1M images in MedicalNarratives contains localized mouse traces in image pixels, creating a spatial and temporal association between the text and pixels. To evaluate the utility of MedicalNarratives, we train GenMedClip with a CLIP-like objective using our dataset spanning 12 medical domains. GenMedClip outperforms previous state-of-the-art models on all 12 domains on a newly constructed medical imaging benchmark. $\href{https://huggingface.co/datasets/wisdomik/MedicalNarratives}{[Data]}$
format Preprint
id arxiv_https___arxiv_org_abs_2501_04184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
Ikezogwo, Wisdom O.
Zhang, Kevin
Seyfioglu, Mehmet Saygin
Ghezloo, Fatemeh
Shapiro, Linda
Krishna, Ranjay
Computer Vision and Pattern Recognition
Multi-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury. To enable representation learning for medical images at scale, we turn to YouTube, a platform with a large reservoir of open-source medical pedagogical videos. We curate MedicalNarratives, a dataset 4.7M medical image-text pairs, with 1M samples containing dense annotations in the form of spatial traces (and bounding boxes), and 118K videos centered on the trace event (with aligned text), enabling spatiotemporal grounding beyond single frames. Similar to $\textit{think-aloud}$ studies where instructors speak while hovering their mouse cursor movements over relevant image regions, 1M images in MedicalNarratives contains localized mouse traces in image pixels, creating a spatial and temporal association between the text and pixels. To evaluate the utility of MedicalNarratives, we train GenMedClip with a CLIP-like objective using our dataset spanning 12 medical domains. GenMedClip outperforms previous state-of-the-art models on all 12 domains on a newly constructed medical imaging benchmark. $\href{https://huggingface.co/datasets/wisdomik/MedicalNarratives}{[Data]}$
title MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.04184