PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bahaj, Adil, Fadi, Oumaima, Chetouani, Mohamed, Ghogho, Mounir
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914008336957440
author Bahaj, Adil
Fadi, Oumaima
Chetouani, Mohamed
Ghogho, Mounir
author_facet Bahaj, Adil
Fadi, Oumaima
Chetouani, Mohamed
Ghogho, Mounir
contents Large language models (LLMs) and vision-augmented LLMs (VLMs) have significantly advanced medical informatics, diagnostics, and decision support. However, these models exhibit systematic biases, particularly age bias, compromising their reliability and equity. This is evident in their poorer performance on pediatric-focused text and visual question-answering tasks. This bias reflects a broader imbalance in medical research, where pediatric studies receive less funding and representation despite the significant disease burden in children. To address these issues, a new comprehensive multi-modal pediatric question-answering benchmark, PediatricsMQA, has been introduced. It consists of 3,417 text-based multiple-choice questions (MCQs) covering 131 pediatric topics across seven developmental stages (prenatal to adolescent) and 2,067 vision-based MCQs using 634 pediatric images from 67 imaging modalities and 256 anatomical regions. The dataset was developed using a hybrid manual-automatic pipeline, incorporating peer-reviewed pediatric literature, validated question banks, existing benchmarks, and existing QA resources. Evaluating state-of-the-art open models, we find dramatic performance drops in younger cohorts, highlighting the need for age-aware methods to ensure equitable AI support in pediatric care.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
Bahaj, Adil
Fadi, Oumaima
Chetouani, Mohamed
Ghogho, Mounir
Computers and Society
Artificial Intelligence
Computation and Language
Graphics
Multimedia
Large language models (LLMs) and vision-augmented LLMs (VLMs) have significantly advanced medical informatics, diagnostics, and decision support. However, these models exhibit systematic biases, particularly age bias, compromising their reliability and equity. This is evident in their poorer performance on pediatric-focused text and visual question-answering tasks. This bias reflects a broader imbalance in medical research, where pediatric studies receive less funding and representation despite the significant disease burden in children. To address these issues, a new comprehensive multi-modal pediatric question-answering benchmark, PediatricsMQA, has been introduced. It consists of 3,417 text-based multiple-choice questions (MCQs) covering 131 pediatric topics across seven developmental stages (prenatal to adolescent) and 2,067 vision-based MCQs using 634 pediatric images from 67 imaging modalities and 256 anatomical regions. The dataset was developed using a hybrid manual-automatic pipeline, incorporating peer-reviewed pediatric literature, validated question banks, existing benchmarks, and existing QA resources. Evaluating state-of-the-art open models, we find dramatic performance drops in younger cohorts, highlighting the need for age-aware methods to ensure equitable AI support in pediatric care.
title PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
topic Computers and Society
Artificial Intelligence
Computation and Language
Graphics
Multimedia
url https://arxiv.org/abs/2508.16439