RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Butsanets, Léo, Corbière, Charles, Khlaut, Julien, Manceron, Pierre, Dancette, Corentin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917367815077888
author Butsanets, Léo
Corbière, Charles
Khlaut, Julien
Manceron, Pierre
Dancette, Corentin
author_facet Butsanets, Léo
Corbière, Charles
Khlaut, Julien
Manceron, Pierre
Dancette, Corentin
contents In this work, we introduce RadImageNet-VQA, a large-scale dataset designed to advance radiologic visual question answering (VQA) on CT and MRI exams. Existing medical VQA datasets are limited in scale, dominated by X-ray imaging or biomedical illustrations, and often prone to text-based shortcuts. RadImageNet-VQA is built from expert-curated annotations and provides 750K images paired with 7.5M question-answer samples. It covers three key tasks - abnormality detection, anatomy recognition, and pathology identification - spanning eight anatomical regions and 97 pathology categories, and supports open-ended, closed-ended, and multiple-choice questions. Extensive experiments show that state-of-the-art vision-language models still struggle with fine-grained pathology identification, particularly in open-ended settings and even after fine-tuning. Text-only analysis further reveals that model performance collapses to near-random without image inputs, confirming that RadImageNet-VQA is free from linguistic shortcuts. The full dataset and benchmark are publicly available at https://huggingface.co/datasets/raidium/RadImageNet-VQA.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17396
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
Butsanets, Léo
Corbière, Charles
Khlaut, Julien
Manceron, Pierre
Dancette, Corentin
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
In this work, we introduce RadImageNet-VQA, a large-scale dataset designed to advance radiologic visual question answering (VQA) on CT and MRI exams. Existing medical VQA datasets are limited in scale, dominated by X-ray imaging or biomedical illustrations, and often prone to text-based shortcuts. RadImageNet-VQA is built from expert-curated annotations and provides 750K images paired with 7.5M question-answer samples. It covers three key tasks - abnormality detection, anatomy recognition, and pathology identification - spanning eight anatomical regions and 97 pathology categories, and supports open-ended, closed-ended, and multiple-choice questions. Extensive experiments show that state-of-the-art vision-language models still struggle with fine-grained pathology identification, particularly in open-ended settings and even after fine-tuning. Text-only analysis further reveals that model performance collapses to near-random without image inputs, confirming that RadImageNet-VQA is free from linguistic shortcuts. The full dataset and benchmark are publicly available at https://huggingface.co/datasets/raidium/RadImageNet-VQA.
title RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.17396