Common Task Framework For a Critical Evaluation of Scientific Machine Learning Algorithms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wyder, Philippe Martin, Goldfeder, Judah, Yermakov, Alexey, Zhao, Yue, Riva, Stefano, Williams, Jan P., Zoro, David, Rude, Amy Sara, Tomasetto, Matteo, Germany, Joe, Bakarji, Joseph, Maierhofer, Georg, Cranmer, Miles, Kutz, J. Nathan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914465487781888
author Wyder, Philippe Martin
Goldfeder, Judah
Yermakov, Alexey
Zhao, Yue
Riva, Stefano
Williams, Jan P.
Zoro, David
Rude, Amy Sara
Tomasetto, Matteo
Germany, Joe
Bakarji, Joseph
Maierhofer, Georg
Cranmer, Miles
Kutz, J. Nathan
author_facet Wyder, Philippe Martin
Goldfeder, Judah
Yermakov, Alexey
Zhao, Yue
Riva, Stefano
Williams, Jan P.
Zoro, David
Rude, Amy Sara
Tomasetto, Matteo
Germany, Joe
Bakarji, Joseph
Maierhofer, Georg
Cranmer, Miles
Kutz, J. Nathan
contents Machine learning (ML) is transforming modeling and control in the physical, engineering, and biological sciences. However, rapid development has outpaced the creation of standardized, objective benchmarks - leading to weak baselines, reporting bias, and inconsistent evaluations across methods. This undermines reproducibility, misguides resource allocation, and obscures scientific progress. To address this, we propose a Common Task Framework (CTF) for scientific machine learning. The CTF features a curated set of datasets and task-specific metrics spanning forecasting, state reconstruction, and generalization under realistic constraints, including noise and limited data. Inspired by the success of CTFs in fields like natural language processing and computer vision, our framework provides a structured, rigorous foundation for head-to-head evaluation of diverse algorithms. As a first step, we benchmark methods on two canonical nonlinear systems: Kuramoto-Sivashinsky and Lorenz. These results illustrate the utility of the CTF in revealing method strengths, limitations, and suitability for specific classes of problems and diverse objectives. Next, we are launching a competition around a global real world sea surface temperature dataset with a true holdout dataset to foster community engagement. Our long-term vision is to replace ad hoc comparisons with standardized evaluations on hidden test sets that raise the bar for rigor and reproducibility in scientific ML.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23166
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Common Task Framework For a Critical Evaluation of Scientific Machine Learning Algorithms
Wyder, Philippe Martin
Goldfeder, Judah
Yermakov, Alexey
Zhao, Yue
Riva, Stefano
Williams, Jan P.
Zoro, David
Rude, Amy Sara
Tomasetto, Matteo
Germany, Joe
Bakarji, Joseph
Maierhofer, Georg
Cranmer, Miles
Kutz, J. Nathan
Computational Engineering, Finance, and Science
Computational Physics
Machine learning (ML) is transforming modeling and control in the physical, engineering, and biological sciences. However, rapid development has outpaced the creation of standardized, objective benchmarks - leading to weak baselines, reporting bias, and inconsistent evaluations across methods. This undermines reproducibility, misguides resource allocation, and obscures scientific progress. To address this, we propose a Common Task Framework (CTF) for scientific machine learning. The CTF features a curated set of datasets and task-specific metrics spanning forecasting, state reconstruction, and generalization under realistic constraints, including noise and limited data. Inspired by the success of CTFs in fields like natural language processing and computer vision, our framework provides a structured, rigorous foundation for head-to-head evaluation of diverse algorithms. As a first step, we benchmark methods on two canonical nonlinear systems: Kuramoto-Sivashinsky and Lorenz. These results illustrate the utility of the CTF in revealing method strengths, limitations, and suitability for specific classes of problems and diverse objectives. Next, we are launching a competition around a global real world sea surface temperature dataset with a true holdout dataset to foster community engagement. Our long-term vision is to replace ad hoc comparisons with standardized evaluations on hidden test sets that raise the bar for rigor and reproducibility in scientific ML.
title Common Task Framework For a Critical Evaluation of Scientific Machine Learning Algorithms
topic Computational Engineering, Finance, and Science
Computational Physics
url https://arxiv.org/abs/2510.23166