Towards reliable use of artificial intelligence to classify otitis media using otoscopic images: Addressing bias and improving data quality

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yixi, Habib, Al-Rahim, Crossland, Graeme, Patel, Hemi, Perry, Chris, Bock, Kris, Lian, Tony, Weeks, William B., Dodhia, Rahul, Ferres, Juan Lavista, Singh, Narinder Pal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908487445905408
author Xu, Yixi
Habib, Al-Rahim
Crossland, Graeme
Patel, Hemi
Perry, Chris
Bock, Kris
Lian, Tony
Weeks, William B.
Dodhia, Rahul
Ferres, Juan Lavista
Singh, Narinder Pal
author_facet Xu, Yixi
Habib, Al-Rahim
Crossland, Graeme
Patel, Hemi
Perry, Chris
Bock, Kris
Lian, Tony
Weeks, William B.
Dodhia, Rahul
Ferres, Juan Lavista
Singh, Narinder Pal
contents Ear disease contributes significantly to global hearing loss, with recurrent otitis media being a primary preventable cause in children, impacting development. Artificial intelligence (AI) offers promise for early diagnosis via otoscopic image analysis, but dataset biases and inconsistencies limit model generalizability and reliability. This retrospective study systematically evaluated three public otoscopic image datasets (Chile; Ohio, USA; Türkiye) using quantitative and qualitative methods. Two counterfactual experiments were performed: (1) obscuring clinically relevant features to assess model reliance on non-clinical artifacts, and (2) evaluating the impact of hue, saturation, and value on diagnostic outcomes. Quantitative analysis revealed significant biases in the Chile and Ohio, USA datasets. Counterfactual Experiment I found high internal performance (AUC > 0.90) but poor external generalization, because of dataset-specific artifacts. The Türkiye dataset had fewer biases, with AUC decreasing from 0.86 to 0.65 as masking increased, suggesting higher reliance on clinically meaningful features. Counterfactual Experiment II identified common artifacts in the Chile and Ohio, USA datasets. A logistic regression model trained on clinically irrelevant features from the Chile dataset achieved high internal (AUC = 0.89) and external (Ohio, USA: AUC = 0.87) performance. Qualitative analysis identified redundancy in all the datasets and stylistic biases in the Ohio, USA dataset that correlated with clinical outcomes. In summary, dataset biases significantly compromise reliability and generalizability of AI-based otoscopic diagnostic models. Addressing these biases through standardized imaging protocols, diverse dataset inclusion, and improved labeling methods is crucial for developing robust AI solutions, improving high-quality healthcare access, and enhancing diagnostic accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18842
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards reliable use of artificial intelligence to classify otitis media using otoscopic images: Addressing bias and improving data quality
Xu, Yixi
Habib, Al-Rahim
Crossland, Graeme
Patel, Hemi
Perry, Chris
Bock, Kris
Lian, Tony
Weeks, William B.
Dodhia, Rahul
Ferres, Juan Lavista
Singh, Narinder Pal
Computers and Society
Ear disease contributes significantly to global hearing loss, with recurrent otitis media being a primary preventable cause in children, impacting development. Artificial intelligence (AI) offers promise for early diagnosis via otoscopic image analysis, but dataset biases and inconsistencies limit model generalizability and reliability. This retrospective study systematically evaluated three public otoscopic image datasets (Chile; Ohio, USA; Türkiye) using quantitative and qualitative methods. Two counterfactual experiments were performed: (1) obscuring clinically relevant features to assess model reliance on non-clinical artifacts, and (2) evaluating the impact of hue, saturation, and value on diagnostic outcomes. Quantitative analysis revealed significant biases in the Chile and Ohio, USA datasets. Counterfactual Experiment I found high internal performance (AUC > 0.90) but poor external generalization, because of dataset-specific artifacts. The Türkiye dataset had fewer biases, with AUC decreasing from 0.86 to 0.65 as masking increased, suggesting higher reliance on clinically meaningful features. Counterfactual Experiment II identified common artifacts in the Chile and Ohio, USA datasets. A logistic regression model trained on clinically irrelevant features from the Chile dataset achieved high internal (AUC = 0.89) and external (Ohio, USA: AUC = 0.87) performance. Qualitative analysis identified redundancy in all the datasets and stylistic biases in the Ohio, USA dataset that correlated with clinical outcomes. In summary, dataset biases significantly compromise reliability and generalizability of AI-based otoscopic diagnostic models. Addressing these biases through standardized imaging protocols, diverse dataset inclusion, and improved labeling methods is crucial for developing robust AI solutions, improving high-quality healthcare access, and enhancing diagnostic accuracy.
title Towards reliable use of artificial intelligence to classify otitis media using otoscopic images: Addressing bias and improving data quality
topic Computers and Society
url https://arxiv.org/abs/2507.18842