Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Subramanian, Shashank, Harrington, Peter, Keutzer, Kurt, Bhimji, Wahid, Morozov, Dmitriy, Mahoney, Michael, Gholami, Amir
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909060540923904
author Subramanian, Shashank
Harrington, Peter
Keutzer, Kurt
Bhimji, Wahid
Morozov, Dmitriy
Mahoney, Michael
Gholami, Amir
author_facet Subramanian, Shashank
Harrington, Peter
Keutzer, Kurt
Bhimji, Wahid
Morozov, Dmitriy
Mahoney, Michael
Gholami, Amir
contents Pre-trained machine learning (ML) models have shown great performance for a wide range of applications, in particular in natural language processing (NLP) and computer vision (CV). Here, we study how pre-training could be used for scientific machine learning (SciML) applications, specifically in the context of transfer learning. We study the transfer behavior of these models as (i) the pre-trained model size is scaled, (ii) the downstream training dataset size is scaled, (iii) the physics parameters are systematically pushed out of distribution, and (iv) how a single model pre-trained on a mixture of different physics problems can be adapted to various downstream applications. We find that-when fine-tuned appropriately-transfer learning can help reach desired accuracy levels with orders of magnitude fewer downstream examples (across different tasks that can even be out-of-distribution) than training from scratch, with consistent behavior across a wide range of downstream examples. We also find that fine-tuning these models yields more performance gains as model size increases, compared to training from scratch on new downstream tasks. These results hold for a broad range of PDE learning tasks. All in all, our results demonstrate the potential of the "pre-train and fine-tune" paradigm for SciML problems, demonstrating a path towards building SciML foundation models. We open-source our code for reproducibility.
format Preprint
id arxiv_https___arxiv_org_abs_2306_00258
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
Subramanian, Shashank
Harrington, Peter
Keutzer, Kurt
Bhimji, Wahid
Morozov, Dmitriy
Mahoney, Michael
Gholami, Amir
Machine Learning
Numerical Analysis
Pre-trained machine learning (ML) models have shown great performance for a wide range of applications, in particular in natural language processing (NLP) and computer vision (CV). Here, we study how pre-training could be used for scientific machine learning (SciML) applications, specifically in the context of transfer learning. We study the transfer behavior of these models as (i) the pre-trained model size is scaled, (ii) the downstream training dataset size is scaled, (iii) the physics parameters are systematically pushed out of distribution, and (iv) how a single model pre-trained on a mixture of different physics problems can be adapted to various downstream applications. We find that-when fine-tuned appropriately-transfer learning can help reach desired accuracy levels with orders of magnitude fewer downstream examples (across different tasks that can even be out-of-distribution) than training from scratch, with consistent behavior across a wide range of downstream examples. We also find that fine-tuning these models yields more performance gains as model size increases, compared to training from scratch on new downstream tasks. These results hold for a broad range of PDE learning tasks. All in all, our results demonstrate the potential of the "pre-train and fine-tune" paradigm for SciML problems, demonstrating a path towards building SciML foundation models. We open-source our code for reproducibility.
title Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
topic Machine Learning
Numerical Analysis
url https://arxiv.org/abs/2306.00258