BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-Rays

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yang, Faith, Tan Li Hui, Xu, Yanyu, Leng, Sicong, Xu, Xinxing, Liu, Yong, Goh, Rick Siow Mong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914996895612928
author Zhou, Yang
Faith, Tan Li Hui
Xu, Yanyu
Leng, Sicong
Xu, Xinxing
Liu, Yong
Goh, Rick Siow Mong
author_facet Zhou, Yang
Faith, Tan Li Hui
Xu, Yanyu
Leng, Sicong
Xu, Xinxing
Liu, Yong
Goh, Rick Siow Mong
contents Medical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. MedVLP can provide useful features to downstream tasks and facilitate adapting task-specific models to new setups using fewer examples. However, existing MedVLP methods often differ in terms of datasets, preprocessing, and finetuning implementations. This pose great challenges in evaluating how well a MedVLP method generalizes to various clinically-relevant tasks due to the lack of unified, standardized, and comprehensive benchmark. To fill this gap, we propose BenchX, a unified benchmark framework that enables head-to-head comparison and systematical analysis between MedVLP methods using public chest X-ray datasets. Specifically, BenchX is composed of three components: 1) Comprehensive datasets covering nine datasets and four medical tasks; 2) Benchmark suites to standardize data preprocessing, train-test splits, and parameter selection; 3) Unified finetuning protocols that accommodate heterogeneous MedVLP methods for consistent task adaptation in classification, segmentation, and report generation, respectively. Utilizing BenchX, we establish baselines for nine state-of-the-art MedVLP methods and found that the performance of some early MedVLP methods can be enhanced to surpass more recent ones, prompting a revisiting of the developments and conclusions from prior works in MedVLP. Our code are available at https://github.com/yangzhou12/BenchX.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21969
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-Rays
Zhou, Yang
Faith, Tan Li Hui
Xu, Yanyu
Leng, Sicong
Xu, Xinxing
Liu, Yong
Goh, Rick Siow Mong
Computer Vision and Pattern Recognition
Medical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. MedVLP can provide useful features to downstream tasks and facilitate adapting task-specific models to new setups using fewer examples. However, existing MedVLP methods often differ in terms of datasets, preprocessing, and finetuning implementations. This pose great challenges in evaluating how well a MedVLP method generalizes to various clinically-relevant tasks due to the lack of unified, standardized, and comprehensive benchmark. To fill this gap, we propose BenchX, a unified benchmark framework that enables head-to-head comparison and systematical analysis between MedVLP methods using public chest X-ray datasets. Specifically, BenchX is composed of three components: 1) Comprehensive datasets covering nine datasets and four medical tasks; 2) Benchmark suites to standardize data preprocessing, train-test splits, and parameter selection; 3) Unified finetuning protocols that accommodate heterogeneous MedVLP methods for consistent task adaptation in classification, segmentation, and report generation, respectively. Utilizing BenchX, we establish baselines for nine state-of-the-art MedVLP methods and found that the performance of some early MedVLP methods can be enhanced to surpass more recent ones, prompting a revisiting of the developments and conclusions from prior works in MedVLP. Our code are available at https://github.com/yangzhou12/BenchX.
title BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-Rays
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.21969