CT-Bench: A Benchmark for Multimodal Lesion Understanding in Computed Tomography

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Qingqing, Jin, Qiao, Mathai, Tejas S., Fang, Yin, Wang, Zhizheng, Yang, Yifan, Sarfo-Gyamfi, Maame, Hou, Benjamin, Gu, Ran, Balamuralikrishna, Praveen T. S., Wang, Kenneth C., Summers, Ronald M., Lu, Zhiyong
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917282598354944
author Zhu, Qingqing
Jin, Qiao
Mathai, Tejas S.
Fang, Yin
Wang, Zhizheng
Yang, Yifan
Sarfo-Gyamfi, Maame
Hou, Benjamin
Gu, Ran
Balamuralikrishna, Praveen T. S.
Wang, Kenneth C.
Summers, Ronald M.
Lu, Zhiyong
author_facet Zhu, Qingqing
Jin, Qiao
Mathai, Tejas S.
Fang, Yin
Wang, Zhizheng
Yang, Yifan
Sarfo-Gyamfi, Maame
Hou, Benjamin
Gu, Ran
Balamuralikrishna, Praveen T. S.
Wang, Kenneth C.
Summers, Ronald M.
Lu, Zhiyong
contents Artificial intelligence (AI) can automatically delineate lesions on computed tomography (CT) and generate radiology report content, yet progress is limited by the scarcity of publicly available CT datasets with lesion-level annotations. To bridge this gap, we introduce CT-Bench, a first-of-its-kind benchmark dataset comprising two components: a Lesion Image and Metadata Set containing 20,335 lesions from 7,795 CT studies with bounding boxes, descriptions, and size information, and a multitask visual question answering benchmark with 2,850 QA pairs covering lesion localization, description, size estimation, and attribute categorization. Hard negative examples are included to reflect real-world diagnostic challenges. We evaluate multiple state-of-the-art multimodal models, including vision-language and medical CLIP variants, by comparing their performance to radiologist assessments, demonstrating the value of CT-Bench as a comprehensive benchmark for lesion analysis. Moreover, fine-tuning models on the Lesion Image and Metadata Set yields significant performance gains across both components, underscoring the clinical utility of CT-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14879
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CT-Bench: A Benchmark for Multimodal Lesion Understanding in Computed Tomography
Zhu, Qingqing
Jin, Qiao
Mathai, Tejas S.
Fang, Yin
Wang, Zhizheng
Yang, Yifan
Sarfo-Gyamfi, Maame
Hou, Benjamin
Gu, Ran
Balamuralikrishna, Praveen T. S.
Wang, Kenneth C.
Summers, Ronald M.
Lu, Zhiyong
Computer Vision and Pattern Recognition
Artificial Intelligence
Artificial intelligence (AI) can automatically delineate lesions on computed tomography (CT) and generate radiology report content, yet progress is limited by the scarcity of publicly available CT datasets with lesion-level annotations. To bridge this gap, we introduce CT-Bench, a first-of-its-kind benchmark dataset comprising two components: a Lesion Image and Metadata Set containing 20,335 lesions from 7,795 CT studies with bounding boxes, descriptions, and size information, and a multitask visual question answering benchmark with 2,850 QA pairs covering lesion localization, description, size estimation, and attribute categorization. Hard negative examples are included to reflect real-world diagnostic challenges. We evaluate multiple state-of-the-art multimodal models, including vision-language and medical CLIP variants, by comparing their performance to radiologist assessments, demonstrating the value of CT-Bench as a comprehensive benchmark for lesion analysis. Moreover, fine-tuning models on the Lesion Image and Metadata Set yields significant performance gains across both components, underscoring the clinical utility of CT-Bench.
title CT-Bench: A Benchmark for Multimodal Lesion Understanding in Computed Tomography
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.14879