Multifidelity linear regression for scientific machine learning from scarce data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Elizabeth, Kang, Dayoung, Sella, Vignesh, Chaudhuri, Anirban
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909237656944640
author Qian, Elizabeth
Kang, Dayoung
Sella, Vignesh
Chaudhuri, Anirban
author_facet Qian, Elizabeth
Kang, Dayoung
Sella, Vignesh
Chaudhuri, Anirban
contents Machine learning (ML) methods, which fit to data the parameters of a given parameterized model class, have garnered significant interest as potential methods for learning surrogate models for complex engineering systems for which traditional simulation is expensive. However, in many scientific and engineering settings, generating high-fidelity data on which to train ML models is expensive, and the available budget for generating training data is limited, so that high-fidelity training data are scarce. ML models trained on scarce data have high variance, resulting in poor expected generalization performance. We propose a new multifidelity training approach for scientific machine learning via linear regression that exploits the scientific context where data of varying fidelities and costs are available: for example, high-fidelity data may be generated by an expensive fully resolved physics simulation whereas lower-fidelity data may arise from a cheaper model based on simplifying assumptions. We use the multifidelity data within an approximate control variate framework to define new multifidelity Monte Carlo estimators for linear regression models. We provide bias and variance analysis of our new estimators that guarantee the approach's accuracy and improved robustness to scarce high-fidelity data. Numerical results demonstrate that our multifidelity training approach achieves similar accuracy to the standard high-fidelity only approach with orders-of-magnitude reduced high-fidelity data requirements.
format Preprint
id arxiv_https___arxiv_org_abs_2403_08627
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multifidelity linear regression for scientific machine learning from scarce data
Qian, Elizabeth
Kang, Dayoung
Sella, Vignesh
Chaudhuri, Anirban
Machine Learning
Computational Engineering, Finance, and Science
Machine learning (ML) methods, which fit to data the parameters of a given parameterized model class, have garnered significant interest as potential methods for learning surrogate models for complex engineering systems for which traditional simulation is expensive. However, in many scientific and engineering settings, generating high-fidelity data on which to train ML models is expensive, and the available budget for generating training data is limited, so that high-fidelity training data are scarce. ML models trained on scarce data have high variance, resulting in poor expected generalization performance. We propose a new multifidelity training approach for scientific machine learning via linear regression that exploits the scientific context where data of varying fidelities and costs are available: for example, high-fidelity data may be generated by an expensive fully resolved physics simulation whereas lower-fidelity data may arise from a cheaper model based on simplifying assumptions. We use the multifidelity data within an approximate control variate framework to define new multifidelity Monte Carlo estimators for linear regression models. We provide bias and variance analysis of our new estimators that guarantee the approach's accuracy and improved robustness to scarce high-fidelity data. Numerical results demonstrate that our multifidelity training approach achieves similar accuracy to the standard high-fidelity only approach with orders-of-magnitude reduced high-fidelity data requirements.
title Multifidelity linear regression for scientific machine learning from scarce data
topic Machine Learning
Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2403.08627