How many samples to label for an application given a foundation model? Chest X-ray classification study

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nechaev, Nikolay, Przhezdzetskaia, Evgeniia, Gombolevskiy, Viktor, Umerenkov, Dmitry, Dylov, Dmitry
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914107435778048
author Nechaev, Nikolay
Przhezdzetskaia, Evgeniia
Gombolevskiy, Viktor
Umerenkov, Dmitry
Dylov, Dmitry
author_facet Nechaev, Nikolay
Przhezdzetskaia, Evgeniia
Gombolevskiy, Viktor
Umerenkov, Dmitry
Dylov, Dmitry
contents Chest X-ray classification is vital yet resource-intensive, typically demanding extensive annotated data for accurate diagnosis. Foundation models mitigate this reliance, but how many labeled samples are required remains unclear. We systematically evaluate the use of power-law fits to predict the training size necessary for specific ROC-AUC thresholds. Testing multiple pathologies and foundation models, we find XrayCLIP and XraySigLIP achieve strong performance with significantly fewer labeled examples than a ResNet-50 baseline. Importantly, learning curve slopes from just 50 labeled cases accurately forecast final performance plateaus. Our results enable practitioners to minimize annotation costs by labeling only the essential samples for targeted performance.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11553
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How many samples to label for an application given a foundation model? Chest X-ray classification study
Nechaev, Nikolay
Przhezdzetskaia, Evgeniia
Gombolevskiy, Viktor
Umerenkov, Dmitry
Dylov, Dmitry
Computer Vision and Pattern Recognition
68T07 (Primary) 68T45, 62H30, 62P10 (Secondary)
Chest X-ray classification is vital yet resource-intensive, typically demanding extensive annotated data for accurate diagnosis. Foundation models mitigate this reliance, but how many labeled samples are required remains unclear. We systematically evaluate the use of power-law fits to predict the training size necessary for specific ROC-AUC thresholds. Testing multiple pathologies and foundation models, we find XrayCLIP and XraySigLIP achieve strong performance with significantly fewer labeled examples than a ResNet-50 baseline. Importantly, learning curve slopes from just 50 labeled cases accurately forecast final performance plateaus. Our results enable practitioners to minimize annotation costs by labeling only the essential samples for targeted performance.
title How many samples to label for an application given a foundation model? Chest X-ray classification study
topic Computer Vision and Pattern Recognition
68T07 (Primary) 68T45, 62H30, 62P10 (Secondary)
url https://arxiv.org/abs/2510.11553