Benchmarking Data Efficiency in $Δ$-ML and Multifidelity Models for Quantum Chemistry

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Vinod, Vivin, Zaspel, Peter
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910890922606592
author Vinod, Vivin
Zaspel, Peter
author_facet Vinod, Vivin
Zaspel, Peter
contents The development of machine learning (ML) methods has made quantum chemistry (QC) calculations more accessible by reducing the compute cost incurred in conventional QC methods. This has since been translated into the overhead cost of generating training data. Increased work in reducing the cost of generating training data resulted in the development of $Δ$-ML and multifidelity machine learning methods which use data at more than one QC level of accuracy, or fidelity. This work compares the data costs associated with $Δ$-ML, multifidelity machine learning (MFML), and optimized MFML (o-MFML) in contrast with a newly introduced Multifidelity$Δ$-Machine Learning (MF$Δ$ML) method for the prediction of ground state energies, vertical excitation energies, and the magnitude of electronic contribution of molecular dipole moments from the multifidelity benchmark dataset QeMFi. This assessment is made on the basis of training data generation cost associated with each model and is compared with the single fidelity kernel ridge regression (KRR) case. The results indicate that the use of multifidelity methods surpasses the standard $Δ$-ML approaches in cases of a large number of predictions. For applications which require only a few evaluations to be made using ML models, while the $Δ$-ML method might be favored, the MF$Δ$ML method is shown to be more efficient.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11391
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Benchmarking Data Efficiency in $Δ$-ML and Multifidelity Models for Quantum Chemistry
Vinod, Vivin
Zaspel, Peter
Chemical Physics
Machine Learning
Computational Physics
The development of machine learning (ML) methods has made quantum chemistry (QC) calculations more accessible by reducing the compute cost incurred in conventional QC methods. This has since been translated into the overhead cost of generating training data. Increased work in reducing the cost of generating training data resulted in the development of $Δ$-ML and multifidelity machine learning methods which use data at more than one QC level of accuracy, or fidelity. This work compares the data costs associated with $Δ$-ML, multifidelity machine learning (MFML), and optimized MFML (o-MFML) in contrast with a newly introduced Multifidelity$Δ$-Machine Learning (MF$Δ$ML) method for the prediction of ground state energies, vertical excitation energies, and the magnitude of electronic contribution of molecular dipole moments from the multifidelity benchmark dataset QeMFi. This assessment is made on the basis of training data generation cost associated with each model and is compared with the single fidelity kernel ridge regression (KRR) case. The results indicate that the use of multifidelity methods surpasses the standard $Δ$-ML approaches in cases of a large number of predictions. For applications which require only a few evaluations to be made using ML models, while the $Δ$-ML method might be favored, the MF$Δ$ML method is shown to be more efficient.
title Benchmarking Data Efficiency in $Δ$-ML and Multifidelity Models for Quantum Chemistry
topic Chemical Physics
Machine Learning
Computational Physics
url https://arxiv.org/abs/2410.11391