Data Readiness for Scientific AI at Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Brewer, Wesley, Widener, Patrick, Anantharaj, Valentine, Wang, Feiyi, Beck, Tom, Shankar, Arjun, Oral, Sarp
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908473312149504
author Brewer, Wesley
Widener, Patrick
Anantharaj, Valentine
Wang, Feiyi
Beck, Tom
Shankar, Arjun
Oral, Sarp
author_facet Brewer, Wesley
Widener, Patrick
Anantharaj, Valentine
Wang, Feiyi
Beck, Tom
Shankar, Arjun
Oral, Sarp
contents This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains - climate, nuclear fusion, bio/health, and materials - to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework composed of Data Readiness Levels (raw to AI-ready) and Data Processing Stages (ingest to shard), both tailored to high performance computing (HPC) environments. This framework outlines key challenges in transforming scientific data for scalable AI training, emphasizing transformer-based generative models. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23018
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data Readiness for Scientific AI at Scale
Brewer, Wesley
Widener, Patrick
Anantharaj, Valentine
Wang, Feiyi
Beck, Tom
Shankar, Arjun
Oral, Sarp
Artificial Intelligence
Computational Engineering, Finance, and Science
Distributed, Parallel, and Cluster Computing
Machine Learning
I.2.6
This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains - climate, nuclear fusion, bio/health, and materials - to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework composed of Data Readiness Levels (raw to AI-ready) and Data Processing Stages (ingest to shard), both tailored to high performance computing (HPC) environments. This framework outlines key challenges in transforming scientific data for scalable AI training, emphasizing transformer-based generative models. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science.
title Data Readiness for Scientific AI at Scale
topic Artificial Intelligence
Computational Engineering, Finance, and Science
Distributed, Parallel, and Cluster Computing
Machine Learning
I.2.6
url https://arxiv.org/abs/2507.23018