Impact of Data Bias on Machine Learning for Crystal Compound Synthesizability Predictions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Davariashtiyani, Ali, Wang, Busheng, Hajinazar, Samad, Zurek, Eva, Kadkhodaei, Sara
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916516150116352
author Davariashtiyani, Ali
Wang, Busheng
Hajinazar, Samad
Zurek, Eva
Kadkhodaei, Sara
author_facet Davariashtiyani, Ali
Wang, Busheng
Hajinazar, Samad
Zurek, Eva
Kadkhodaei, Sara
contents Machine learning models are susceptible to being misled by biases in training data that emphasize incidental correlations over the intended learning task. In this study, we demonstrate the impact of data bias on the performance of a machine learning model designed to predict the synthesizability likelihood of crystal compounds. The model performs a binary classification on labeled crystal samples. Despite using the same architecture for the machine learning model, we showcase how the model's learning and prediction behavior differs once trained on distinct data. We use two data sets for illustration: a mixed-source data set that integrates experimental and computational crystal samples and a single-source data set consisting of data exclusively from one computational database. We present simple procedures to detect data bias and to evaluate its effect on the model's performance and generalization. This study reveals how inconsistent, unbalanced data can propagate bias, undermining real-world applicability even for advanced machine learning techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17956
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Impact of Data Bias on Machine Learning for Crystal Compound Synthesizability Predictions
Davariashtiyani, Ali
Wang, Busheng
Hajinazar, Samad
Zurek, Eva
Kadkhodaei, Sara
Materials Science
Machine learning models are susceptible to being misled by biases in training data that emphasize incidental correlations over the intended learning task. In this study, we demonstrate the impact of data bias on the performance of a machine learning model designed to predict the synthesizability likelihood of crystal compounds. The model performs a binary classification on labeled crystal samples. Despite using the same architecture for the machine learning model, we showcase how the model's learning and prediction behavior differs once trained on distinct data. We use two data sets for illustration: a mixed-source data set that integrates experimental and computational crystal samples and a single-source data set consisting of data exclusively from one computational database. We present simple procedures to detect data bias and to evaluate its effect on the model's performance and generalization. This study reveals how inconsistent, unbalanced data can propagate bias, undermining real-world applicability even for advanced machine learning techniques.
title Impact of Data Bias on Machine Learning for Crystal Compound Synthesizability Predictions
topic Materials Science
url https://arxiv.org/abs/2406.17956