_version_ 1866915744382451712
author Wei, Yishu
Flanders, Adam E.
Colak, Errol
Mongan, John
Prevedello, Luciano M
Chen, Po-Hao
Lee, Henrique Min Ho
Szarf, Gilberto
Shoji, Hamilton
Sho, Jason
Andriole, Katherine
Cook, Tessa
Adams, Lisa C.
Chu, Linda C.
Chung, Maggie
Brusca-Augello, Geraldine
Deva, Djeven P.
Singh, Navneet
Tijmes, Felipe Sanchez
Alpert, Jeffrey B.
Nguyen, Elsie T.
Torigian, Drew A.
Hanneman, Kate
Groner, Lauren K
Phan, Alexander
Islam, Ali
Callejas, Matias F.
Teles, Gustavo Borges da Silva
Jamal, Faisal
Vazirabad, Maryam
Tejani, Ali
Trivedi, Hari
Kuriki, Paulo
Bhayana, Rajesh
Benishay, Elana T.
Lin, Yi
Peng, Yifan
Shih, George
author_facet Wei, Yishu
Flanders, Adam E.
Colak, Errol
Mongan, John
Prevedello, Luciano M
Chen, Po-Hao
Lee, Henrique Min Ho
Szarf, Gilberto
Shoji, Hamilton
Sho, Jason
Andriole, Katherine
Cook, Tessa
Adams, Lisa C.
Chu, Linda C.
Chung, Maggie
Brusca-Augello, Geraldine
Deva, Djeven P.
Singh, Navneet
Tijmes, Felipe Sanchez
Alpert, Jeffrey B.
Nguyen, Elsie T.
Torigian, Drew A.
Hanneman, Kate
Groner, Lauren K
Phan, Alexander
Islam, Ali
Callejas, Matias F.
Teles, Gustavo Borges da Silva
Jamal, Faisal
Vazirabad, Maryam
Tejani, Ali
Trivedi, Hari
Kuriki, Paulo
Bhayana, Rajesh
Benishay, Elana T.
Lin, Yi
Peng, Yifan
Shih, George
contents Multimodal large language models have demonstrated comparable performance to that of radiology trainees on multiple-choice board-style exams. However, to develop clinically useful multimodal LLM tools, high-quality benchmarks curated by domain experts are essential. To curate released and holdout datasets of 100 chest radiographic studies each and propose an artificial intelligence (AI)-assisted expert labeling procedure to allow radiologists to label studies more efficiently. A total of 13,735 deidentified chest radiographs and their corresponding reports from the MIDRC were used. GPT-4o extracted abnormal findings from the reports, which were then mapped to 12 benchmark labels with a locally hosted LLM (Phi-4-Reasoning). From these studies, 1,000 were sampled on the basis of the AI-suggested benchmark labels for expert review; the sampling algorithm ensured that the selected studies were clinically relevant and captured a range of difficulty levels. Seventeen chest radiologists participated, and they marked "Agree all", "Agree mostly" or "Disagree" to indicate their assessment of the correctness of the LLM suggested labels. Each chest radiograph was evaluated by three experts. Of these, at least two radiologists selected "Agree All" for 381 radiographs. From this set, 200 were selected, prioritizing those with less common or multiple finding labels, and divided into 100 released radiographs and 100 reserved as the holdout dataset. The holdout dataset is used exclusively by RSNA to independently evaluate different models. A benchmark of 200 chest radiographic studies with 12 benchmark labels was created and made publicly available https://imaging.rsna.org, with each chest radiograph verified by three radiologists. In addition, an AI-assisted labeling procedure was developed to help radiologists label at scale, minimize unnecessary omissions, and support a semicollaborative environment.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15129
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RSNA Large Language Model Benchmark Dataset for Chest Radiographs of Cardiothoracic Disease: Radiologist Evaluation and Validation Enhanced by AI Labels (REVEAL-CXR)
Wei, Yishu
Flanders, Adam E.
Colak, Errol
Mongan, John
Prevedello, Luciano M
Chen, Po-Hao
Lee, Henrique Min Ho
Szarf, Gilberto
Shoji, Hamilton
Sho, Jason
Andriole, Katherine
Cook, Tessa
Adams, Lisa C.
Chu, Linda C.
Chung, Maggie
Brusca-Augello, Geraldine
Deva, Djeven P.
Singh, Navneet
Tijmes, Felipe Sanchez
Alpert, Jeffrey B.
Nguyen, Elsie T.
Torigian, Drew A.
Hanneman, Kate
Groner, Lauren K
Phan, Alexander
Islam, Ali
Callejas, Matias F.
Teles, Gustavo Borges da Silva
Jamal, Faisal
Vazirabad, Maryam
Tejani, Ali
Trivedi, Hari
Kuriki, Paulo
Bhayana, Rajesh
Benishay, Elana T.
Lin, Yi
Peng, Yifan
Shih, George
Computation and Language
Multimodal large language models have demonstrated comparable performance to that of radiology trainees on multiple-choice board-style exams. However, to develop clinically useful multimodal LLM tools, high-quality benchmarks curated by domain experts are essential. To curate released and holdout datasets of 100 chest radiographic studies each and propose an artificial intelligence (AI)-assisted expert labeling procedure to allow radiologists to label studies more efficiently. A total of 13,735 deidentified chest radiographs and their corresponding reports from the MIDRC were used. GPT-4o extracted abnormal findings from the reports, which were then mapped to 12 benchmark labels with a locally hosted LLM (Phi-4-Reasoning). From these studies, 1,000 were sampled on the basis of the AI-suggested benchmark labels for expert review; the sampling algorithm ensured that the selected studies were clinically relevant and captured a range of difficulty levels. Seventeen chest radiologists participated, and they marked "Agree all", "Agree mostly" or "Disagree" to indicate their assessment of the correctness of the LLM suggested labels. Each chest radiograph was evaluated by three experts. Of these, at least two radiologists selected "Agree All" for 381 radiographs. From this set, 200 were selected, prioritizing those with less common or multiple finding labels, and divided into 100 released radiographs and 100 reserved as the holdout dataset. The holdout dataset is used exclusively by RSNA to independently evaluate different models. A benchmark of 200 chest radiographic studies with 12 benchmark labels was created and made publicly available https://imaging.rsna.org, with each chest radiograph verified by three radiologists. In addition, an AI-assisted labeling procedure was developed to help radiologists label at scale, minimize unnecessary omissions, and support a semicollaborative environment.
title RSNA Large Language Model Benchmark Dataset for Chest Radiographs of Cardiothoracic Disease: Radiologist Evaluation and Validation Enhanced by AI Labels (REVEAL-CXR)
topic Computation and Language
url https://arxiv.org/abs/2601.15129