A General Approach for Determining Applicability Domain of Machine Learning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schultz, Lane E., Wang, Yiqi, Jacobs, Ryan, Morgan, Dane
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913752430936064
author Schultz, Lane E.
Wang, Yiqi
Jacobs, Ryan
Morgan, Dane
author_facet Schultz, Lane E.
Wang, Yiqi
Jacobs, Ryan
Morgan, Dane
contents Knowledge of the domain of applicability of a machine learning model is essential to ensuring accurate and reliable model predictions. In this work, we develop a new and general approach of assessing model domain and demonstrate that our approach provides accurate and meaningful domain designation across multiple model types and material property data sets. Our approach assesses the distance between data in feature space using kernel density estimation, where this distance provides an effective tool for domain determination. We show that chemical groups considered unrelated based on chemical knowledge exhibit significant dissimilarities by our measure. We also show that high measures of dissimilarity are associated with poor model performance (i.e., high residual magnitudes) and poor estimates of model uncertainty (i.e., unreliable uncertainty estimation). Automated tools are provided to enable researchers to establish acceptable dissimilarity thresholds to identify whether new predictions of their own machine learning models are in-domain versus out-of-domain.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05143
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A General Approach for Determining Applicability Domain of Machine Learning Models
Schultz, Lane E.
Wang, Yiqi
Jacobs, Ryan
Morgan, Dane
Materials Science
Other Condensed Matter
Machine Learning
Knowledge of the domain of applicability of a machine learning model is essential to ensuring accurate and reliable model predictions. In this work, we develop a new and general approach of assessing model domain and demonstrate that our approach provides accurate and meaningful domain designation across multiple model types and material property data sets. Our approach assesses the distance between data in feature space using kernel density estimation, where this distance provides an effective tool for domain determination. We show that chemical groups considered unrelated based on chemical knowledge exhibit significant dissimilarities by our measure. We also show that high measures of dissimilarity are associated with poor model performance (i.e., high residual magnitudes) and poor estimates of model uncertainty (i.e., unreliable uncertainty estimation). Automated tools are provided to enable researchers to establish acceptable dissimilarity thresholds to identify whether new predictions of their own machine learning models are in-domain versus out-of-domain.
title A General Approach for Determining Applicability Domain of Machine Learning Models
topic Materials Science
Other Condensed Matter
Machine Learning
url https://arxiv.org/abs/2406.05143