An analysis of data variation and bias in image-based dermatological datasets for machine learning classification

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Filho, Francisco, Santos, Emanoel, Mota, Rodrigo, Cunha, Kelvin, Papais, Fabio, Arruda, Amanda, Baltazar, Mateus, Vieira, Camila, Tavares, José Gabriel, Barros, Rafael, Souza, Othon, Bezerra, Thales, Lopes, Natalia, Moutinho, Érico, Guido, Jéssica, Cruz, Shirley, Borba, Paulo, Ren, Tsang Ing
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912227999612928
author Filho, Francisco
Santos, Emanoel
Mota, Rodrigo
Cunha, Kelvin
Papais, Fabio
Arruda, Amanda
Baltazar, Mateus
Vieira, Camila
Tavares, José Gabriel
Barros, Rafael
Souza, Othon
Bezerra, Thales
Lopes, Natalia
Moutinho, Érico
Guido, Jéssica
Cruz, Shirley
Borba, Paulo
Ren, Tsang Ing
author_facet Filho, Francisco
Santos, Emanoel
Mota, Rodrigo
Cunha, Kelvin
Papais, Fabio
Arruda, Amanda
Baltazar, Mateus
Vieira, Camila
Tavares, José Gabriel
Barros, Rafael
Souza, Othon
Bezerra, Thales
Lopes, Natalia
Moutinho, Érico
Guido, Jéssica
Cruz, Shirley
Borba, Paulo
Ren, Tsang Ing
contents AI algorithms have become valuable in aiding professionals in healthcare. The increasing confidence obtained by these models is helpful in critical decision demands. In clinical dermatology, classification models can detect malignant lesions on patients' skin using only RGB images as input. However, most learning-based methods employ data acquired from dermoscopic datasets on training, which are large and validated by a gold standard. Clinical models aim to deal with classification on users' smartphone cameras that do not contain the corresponding resolution provided by dermoscopy. Also, clinical applications bring new challenges. It can contain captures from uncontrolled environments, skin tone variations, viewpoint changes, noises in data and labels, and unbalanced classes. A possible alternative would be to use transfer learning to deal with the clinical images. However, as the number of samples is low, it can cause degradations on the model's performance; the source distribution used in training differs from the test set. This work aims to evaluate the gap between dermoscopic and clinical samples and understand how the dataset variations impact training. It assesses the main differences between distributions that disturb the model's prediction. Finally, from experiments on different architectures, we argue how to combine the data from divergent distributions, decreasing the impact on the model's final accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2501_08962
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An analysis of data variation and bias in image-based dermatological datasets for machine learning classification
Filho, Francisco
Santos, Emanoel
Mota, Rodrigo
Cunha, Kelvin
Papais, Fabio
Arruda, Amanda
Baltazar, Mateus
Vieira, Camila
Tavares, José Gabriel
Barros, Rafael
Souza, Othon
Bezerra, Thales
Lopes, Natalia
Moutinho, Érico
Guido, Jéssica
Cruz, Shirley
Borba, Paulo
Ren, Tsang Ing
Computer Vision and Pattern Recognition
Artificial Intelligence
I.5.4; J.3
AI algorithms have become valuable in aiding professionals in healthcare. The increasing confidence obtained by these models is helpful in critical decision demands. In clinical dermatology, classification models can detect malignant lesions on patients' skin using only RGB images as input. However, most learning-based methods employ data acquired from dermoscopic datasets on training, which are large and validated by a gold standard. Clinical models aim to deal with classification on users' smartphone cameras that do not contain the corresponding resolution provided by dermoscopy. Also, clinical applications bring new challenges. It can contain captures from uncontrolled environments, skin tone variations, viewpoint changes, noises in data and labels, and unbalanced classes. A possible alternative would be to use transfer learning to deal with the clinical images. However, as the number of samples is low, it can cause degradations on the model's performance; the source distribution used in training differs from the test set. This work aims to evaluate the gap between dermoscopic and clinical samples and understand how the dataset variations impact training. It assesses the main differences between distributions that disturb the model's prediction. Finally, from experiments on different architectures, we argue how to combine the data from divergent distributions, decreasing the impact on the model's final accuracy.
title An analysis of data variation and bias in image-based dermatological datasets for machine learning classification
topic Computer Vision and Pattern Recognition
Artificial Intelligence
I.5.4; J.3
url https://arxiv.org/abs/2501.08962