From Embeddings to Accuracy: Comparing Foundation Models for Radiographic Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xue, Merkow, Jameson, Codella, Noel C. F., Santamaria-Pang, Alberto, Sangani, Naiteek, Ersoy, Alexander, Burt, Christopher, Garrett, John W., Bruce, Richard J., Warner, Joshua D., Bradshaw, Tyler, Tarapov, Ivan, Lungren, Matthew P., McMillan, Alan B.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916931558178816
author Li, Xue
Merkow, Jameson
Codella, Noel C. F.
Santamaria-Pang, Alberto
Sangani, Naiteek
Ersoy, Alexander
Burt, Christopher
Garrett, John W.
Bruce, Richard J.
Warner, Joshua D.
Bradshaw, Tyler
Tarapov, Ivan
Lungren, Matthew P.
McMillan, Alan B.
author_facet Li, Xue
Merkow, Jameson
Codella, Noel C. F.
Santamaria-Pang, Alberto
Sangani, Naiteek
Ersoy, Alexander
Burt, Christopher
Garrett, John W.
Bruce, Richard J.
Warner, Joshua D.
Bradshaw, Tyler
Tarapov, Ivan
Lungren, Matthew P.
McMillan, Alan B.
contents Foundation models provide robust embeddings for diverse tasks, including medical imaging. We evaluate embeddings from seven general and medical-specific foundation models (e.g., DenseNet121, BiomedCLIP, MedImageInsight, Rad-DINO, CXR-Foundation) for training lightweight adapters in multi-class radiography classification. Using a dataset of 8,842 radiographs across seven classes, we trained adapters with algorithms like K-Nearest Neighbors, logistic regression, SVM, random forest, and MLP. The combination of MedImageInsight embeddings with an SVM or MLP adapter achieved the highest mean area under the curve (mAUC) of 93.1%. This performance was statistically superior to other models, including MedSigLIP with an MLP (91.0%), Rad-DINO with an SVM (90.7%), and CXR-Foundation with logistic regression (88.6%). In contrast, models like BiomedCLIP (82.8%) and Med-Flamingo (78.5%) showed lower performance. Crucially, these lightweight adapters are computationally efficient, training in minutes and performing inference in seconds on a CPU, making them practical for clinical use. A fairness analysis of the top-performing MedImageInsight adapter revealed minimal performance disparities across patient gender (within 1.8%) and age groups (std. dev < 1.4%), with no significant statistical differences. These findings confirm that embeddings from specialized foundation models, particularly MedImageInsight, can power accurate, efficient, and equitable diagnostic tools using simple, lightweight adapters.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10823
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Embeddings to Accuracy: Comparing Foundation Models for Radiographic Classification
Li, Xue
Merkow, Jameson
Codella, Noel C. F.
Santamaria-Pang, Alberto
Sangani, Naiteek
Ersoy, Alexander
Burt, Christopher
Garrett, John W.
Bruce, Richard J.
Warner, Joshua D.
Bradshaw, Tyler
Tarapov, Ivan
Lungren, Matthew P.
McMillan, Alan B.
Computer Vision and Pattern Recognition
Image and Video Processing
Foundation models provide robust embeddings for diverse tasks, including medical imaging. We evaluate embeddings from seven general and medical-specific foundation models (e.g., DenseNet121, BiomedCLIP, MedImageInsight, Rad-DINO, CXR-Foundation) for training lightweight adapters in multi-class radiography classification. Using a dataset of 8,842 radiographs across seven classes, we trained adapters with algorithms like K-Nearest Neighbors, logistic regression, SVM, random forest, and MLP. The combination of MedImageInsight embeddings with an SVM or MLP adapter achieved the highest mean area under the curve (mAUC) of 93.1%. This performance was statistically superior to other models, including MedSigLIP with an MLP (91.0%), Rad-DINO with an SVM (90.7%), and CXR-Foundation with logistic regression (88.6%). In contrast, models like BiomedCLIP (82.8%) and Med-Flamingo (78.5%) showed lower performance. Crucially, these lightweight adapters are computationally efficient, training in minutes and performing inference in seconds on a CPU, making them practical for clinical use. A fairness analysis of the top-performing MedImageInsight adapter revealed minimal performance disparities across patient gender (within 1.8%) and age groups (std. dev < 1.4%), with no significant statistical differences. These findings confirm that embeddings from specialized foundation models, particularly MedImageInsight, can power accurate, efficient, and equitable diagnostic tools using simple, lightweight adapters.
title From Embeddings to Accuracy: Comparing Foundation Models for Radiographic Classification
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2505.10823