Common Task Framework For a Critical Evaluation of Scientific Machine Learning Algorithms
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914465487781888 |
|---|---|
| author | Wyder, Philippe Martin Goldfeder, Judah Yermakov, Alexey Zhao, Yue Riva, Stefano Williams, Jan P. Zoro, David Rude, Amy Sara Tomasetto, Matteo Germany, Joe Bakarji, Joseph Maierhofer, Georg Cranmer, Miles Kutz, J. Nathan |
| author_facet | Wyder, Philippe Martin Goldfeder, Judah Yermakov, Alexey Zhao, Yue Riva, Stefano Williams, Jan P. Zoro, David Rude, Amy Sara Tomasetto, Matteo Germany, Joe Bakarji, Joseph Maierhofer, Georg Cranmer, Miles Kutz, J. Nathan |
| contents | Machine learning (ML) is transforming modeling and control in the physical, engineering, and biological sciences. However, rapid development has outpaced the creation of standardized, objective benchmarks - leading to weak baselines, reporting bias, and inconsistent evaluations across methods. This undermines reproducibility, misguides resource allocation, and obscures scientific progress. To address this, we propose a Common Task Framework (CTF) for scientific machine learning. The CTF features a curated set of datasets and task-specific metrics spanning forecasting, state reconstruction, and generalization under realistic constraints, including noise and limited data. Inspired by the success of CTFs in fields like natural language processing and computer vision, our framework provides a structured, rigorous foundation for head-to-head evaluation of diverse algorithms. As a first step, we benchmark methods on two canonical nonlinear systems: Kuramoto-Sivashinsky and Lorenz. These results illustrate the utility of the CTF in revealing method strengths, limitations, and suitability for specific classes of problems and diverse objectives. Next, we are launching a competition around a global real world sea surface temperature dataset with a true holdout dataset to foster community engagement. Our long-term vision is to replace ad hoc comparisons with standardized evaluations on hidden test sets that raise the bar for rigor and reproducibility in scientific ML. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_23166 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Common Task Framework For a Critical Evaluation of Scientific Machine Learning Algorithms Wyder, Philippe Martin Goldfeder, Judah Yermakov, Alexey Zhao, Yue Riva, Stefano Williams, Jan P. Zoro, David Rude, Amy Sara Tomasetto, Matteo Germany, Joe Bakarji, Joseph Maierhofer, Georg Cranmer, Miles Kutz, J. Nathan Computational Engineering, Finance, and Science Computational Physics Machine learning (ML) is transforming modeling and control in the physical, engineering, and biological sciences. However, rapid development has outpaced the creation of standardized, objective benchmarks - leading to weak baselines, reporting bias, and inconsistent evaluations across methods. This undermines reproducibility, misguides resource allocation, and obscures scientific progress. To address this, we propose a Common Task Framework (CTF) for scientific machine learning. The CTF features a curated set of datasets and task-specific metrics spanning forecasting, state reconstruction, and generalization under realistic constraints, including noise and limited data. Inspired by the success of CTFs in fields like natural language processing and computer vision, our framework provides a structured, rigorous foundation for head-to-head evaluation of diverse algorithms. As a first step, we benchmark methods on two canonical nonlinear systems: Kuramoto-Sivashinsky and Lorenz. These results illustrate the utility of the CTF in revealing method strengths, limitations, and suitability for specific classes of problems and diverse objectives. Next, we are launching a competition around a global real world sea surface temperature dataset with a true holdout dataset to foster community engagement. Our long-term vision is to replace ad hoc comparisons with standardized evaluations on hidden test sets that raise the bar for rigor and reproducibility in scientific ML. |
| title | Common Task Framework For a Critical Evaluation of Scientific Machine Learning Algorithms |
| topic | Computational Engineering, Finance, and Science Computational Physics |
| url | https://arxiv.org/abs/2510.23166 |