Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.03232 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918274192637952 |
|---|---|
| author | Bose, Kartik Kumar, Abhinandan Soundararajan, Raghuraman Mudgil, Priya Ralmilay, Samonee Dutta, Niharika Singhal, Manphool Kumar, Arun Sen, Saugata Patra, Anurima Ghosh, Priya Das, Abanti Gupta, Amit Verma, Ashish Sudhakaran, Dipin Dhamija, Ekta Unde, Himangi Kumar, Ishan Rangarajan, Krithika Garg, Prerna Sequeira, Rachel Shylendran, Sudhin Yadav, Taruna Pal, Tej Gupta, Pankaj |
| author_facet | Bose, Kartik Kumar, Abhinandan Soundararajan, Raghuraman Mudgil, Priya Ralmilay, Samonee Dutta, Niharika Singhal, Manphool Kumar, Arun Sen, Saugata Patra, Anurima Ghosh, Priya Das, Abanti Gupta, Amit Verma, Ashish Sudhakaran, Dipin Dhamija, Ekta Unde, Himangi Kumar, Ishan Rangarajan, Krithika Garg, Prerna Sequeira, Rachel Shylendran, Sudhin Yadav, Taruna Pal, Tej Gupta, Pankaj |
| contents | Background: Reporting and Data Systems (RADS) standardize radiology risk communication but automated RADS assignment from narrative reports is challenging because of guideline complexity, output-format constraints, and limited benchmarking across RADS frameworks and model sizes. Purpose: To create RXL-RADSet, a radiologist-verified synthetic multi-RADS benchmark, and compare validity and accuracy of open-weight small language models (SLMs) with a proprietary model for RADS assignment. Materials and Methods: RXL-RADSet contains 1,600 synthetic radiology reports across 10 RADS (BI-RADS, CAD-RADS, GB-RADS, LI-RADS, Lung-RADS, NI-RADS, O-RADS, PI-RADS, TI-RADS, VI-RADS) and multiple modalities. Reports were generated by LLMs using scenario plans and simulated radiologist styles and underwent two-stage radiologist verification. We evaluated 41 quantized SLMs (12 families, 0.135-32B parameters) and GPT-5.2 under a fixed guided prompt. Primary endpoints were validity and accuracy; a secondary analysis compared guided versus zero-shot prompting. Results: Under guided prompting GPT-5.2 achieved 99.8% validity and 81.1% accuracy (1,600 predictions). Pooled SLMs (65,600 predictions) achieved 96.8% validity and 61.1% accuracy; top SLMs in the 20-32B range reached ~99% validity and mid-to-high 70% accuracy. Performance scaled with model size (inflection between <1B and >=10B) and declined with RADS complexity primarily due to classification difficulty rather than invalid outputs. Guided prompting improved validity (99.2% vs 96.7%) and accuracy (78.5% vs 69.6%) compared with zero-shot. Conclusion: RXL-RADSet provides a radiologist-verified multi-RADS benchmark; large SLMs (20-32B) can approach proprietary-model performance under guided prompting, but gaps remain for higher-complexity schemes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_03232 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Multi-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models Bose, Kartik Kumar, Abhinandan Soundararajan, Raghuraman Mudgil, Priya Ralmilay, Samonee Dutta, Niharika Singhal, Manphool Kumar, Arun Sen, Saugata Patra, Anurima Ghosh, Priya Das, Abanti Gupta, Amit Verma, Ashish Sudhakaran, Dipin Dhamija, Ekta Unde, Himangi Kumar, Ishan Rangarajan, Krithika Garg, Prerna Sequeira, Rachel Shylendran, Sudhin Yadav, Taruna Pal, Tej Gupta, Pankaj Computation and Language Artificial Intelligence Background: Reporting and Data Systems (RADS) standardize radiology risk communication but automated RADS assignment from narrative reports is challenging because of guideline complexity, output-format constraints, and limited benchmarking across RADS frameworks and model sizes. Purpose: To create RXL-RADSet, a radiologist-verified synthetic multi-RADS benchmark, and compare validity and accuracy of open-weight small language models (SLMs) with a proprietary model for RADS assignment. Materials and Methods: RXL-RADSet contains 1,600 synthetic radiology reports across 10 RADS (BI-RADS, CAD-RADS, GB-RADS, LI-RADS, Lung-RADS, NI-RADS, O-RADS, PI-RADS, TI-RADS, VI-RADS) and multiple modalities. Reports were generated by LLMs using scenario plans and simulated radiologist styles and underwent two-stage radiologist verification. We evaluated 41 quantized SLMs (12 families, 0.135-32B parameters) and GPT-5.2 under a fixed guided prompt. Primary endpoints were validity and accuracy; a secondary analysis compared guided versus zero-shot prompting. Results: Under guided prompting GPT-5.2 achieved 99.8% validity and 81.1% accuracy (1,600 predictions). Pooled SLMs (65,600 predictions) achieved 96.8% validity and 61.1% accuracy; top SLMs in the 20-32B range reached ~99% validity and mid-to-high 70% accuracy. Performance scaled with model size (inflection between <1B and >=10B) and declined with RADS complexity primarily due to classification difficulty rather than invalid outputs. Guided prompting improved validity (99.2% vs 96.7%) and accuracy (78.5% vs 69.6%) compared with zero-shot. Conclusion: RXL-RADSet provides a radiologist-verified multi-RADS benchmark; large SLMs (20-32B) can approach proprietary-model performance under guided prompting, but gaps remain for higher-complexity schemes. |
| title | Multi-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2601.03232 |