Can we trust deep learning models diagnosis? The impact of domain shift in chest radiograph classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pooch, Eduardo H. P., Ballester, Pedro L., Barros, Rodrigo C.
Format: Preprint
Published: 2019
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916713045426176
author Pooch, Eduardo H. P.
Ballester, Pedro L.
Barros, Rodrigo C.
author_facet Pooch, Eduardo H. P.
Ballester, Pedro L.
Barros, Rodrigo C.
contents While deep learning models become more widespread, their ability to handle unseen data and generalize for any scenario is yet to be challenged. In medical imaging, there is a high heterogeneity of distributions among images based on the equipment that generates them and their parametrization. This heterogeneity triggers a common issue in machine learning called domain shift, which represents the difference between the training data distribution and the distribution of where a model is employed. A high domain shift tends to implicate in a poor generalization performance from the models. In this work, we evaluate the extent of domain shift on four of the largest datasets of chest radiographs. We show how training and testing with different datasets (e.g., training in ChestX-ray14 and testing in CheXpert) drastically affects model performance, posing a big question over the reliability of deep learning models trained on public datasets. We also show that models trained on CheXpert and MIMIC-CXR generalize better to other datasets.
format Preprint
id arxiv_https___arxiv_org_abs_1909_01940
institution arXiv
publishDate 2019
record_format arxiv
spellingShingle Can we trust deep learning models diagnosis? The impact of domain shift in chest radiograph classification
Pooch, Eduardo H. P.
Ballester, Pedro L.
Barros, Rodrigo C.
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
While deep learning models become more widespread, their ability to handle unseen data and generalize for any scenario is yet to be challenged. In medical imaging, there is a high heterogeneity of distributions among images based on the equipment that generates them and their parametrization. This heterogeneity triggers a common issue in machine learning called domain shift, which represents the difference between the training data distribution and the distribution of where a model is employed. A high domain shift tends to implicate in a poor generalization performance from the models. In this work, we evaluate the extent of domain shift on four of the largest datasets of chest radiographs. We show how training and testing with different datasets (e.g., training in ChestX-ray14 and testing in CheXpert) drastically affects model performance, posing a big question over the reliability of deep learning models trained on public datasets. We also show that models trained on CheXpert and MIMIC-CXR generalize better to other datasets.
title Can we trust deep learning models diagnosis? The impact of domain shift in chest radiograph classification
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/1909.01940