What limits performance of weakly supervised deep learning for chest CT classification?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tushar, Fakrul Islam, D'Anniballe, Vincent M., Rubin, Geoffrey D., Lo, Joseph Y.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910320872652800
author Tushar, Fakrul Islam
D'Anniballe, Vincent M.
Rubin, Geoffrey D.
Lo, Joseph Y.
author_facet Tushar, Fakrul Islam
D'Anniballe, Vincent M.
Rubin, Geoffrey D.
Lo, Joseph Y.
contents Weakly supervised learning with noisy data has drawn attention in the medical imaging community due to the sparsity of high-quality disease labels. However, little is known about the limitations of such weakly supervised learning and the effect of these constraints on disease classification performance. In this paper, we test the effects of such weak supervision by examining model tolerance for three conditions. First, we examined model tolerance for noisy data by incrementally increasing error in the labels within the training data. Second, we assessed the impact of dataset size by varying the amount of training data. Third, we compared performance differences between binary and multi-label classification. Results demonstrated that the model could endure up to 10% added label error before experiencing a decline in disease classification performance. Disease classification performance steadily rose as the amount of training data was increased for all disease classes, before experiencing a plateau in performance at 75% of training data. Last, the binary model outperformed the multilabel model in every disease category. However, such interpretations may be misleading, as the binary model was heavily influenced by co-occurring diseases and may not have learned the specific features of the disease in the image. In conclusion, this study may help the medical imaging community understand the benefits and risks of weak supervision with noisy labels. Such studies demonstrate the need to build diverse, large-scale datasets and to develop explainable and responsible AI.
format Preprint
id arxiv_https___arxiv_org_abs_2402_04419
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle What limits performance of weakly supervised deep learning for chest CT classification?
Tushar, Fakrul Islam
D'Anniballe, Vincent M.
Rubin, Geoffrey D.
Lo, Joseph Y.
Image and Video Processing
Machine Learning
Weakly supervised learning with noisy data has drawn attention in the medical imaging community due to the sparsity of high-quality disease labels. However, little is known about the limitations of such weakly supervised learning and the effect of these constraints on disease classification performance. In this paper, we test the effects of such weak supervision by examining model tolerance for three conditions. First, we examined model tolerance for noisy data by incrementally increasing error in the labels within the training data. Second, we assessed the impact of dataset size by varying the amount of training data. Third, we compared performance differences between binary and multi-label classification. Results demonstrated that the model could endure up to 10% added label error before experiencing a decline in disease classification performance. Disease classification performance steadily rose as the amount of training data was increased for all disease classes, before experiencing a plateau in performance at 75% of training data. Last, the binary model outperformed the multilabel model in every disease category. However, such interpretations may be misleading, as the binary model was heavily influenced by co-occurring diseases and may not have learned the specific features of the disease in the image. In conclusion, this study may help the medical imaging community understand the benefits and risks of weak supervision with noisy labels. Such studies demonstrate the need to build diverse, large-scale datasets and to develop explainable and responsible AI.
title What limits performance of weakly supervised deep learning for chest CT classification?
topic Image and Video Processing
Machine Learning
url https://arxiv.org/abs/2402.04419