Data Readiness for Scientific AI at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908473312149504 |
|---|---|
| author | Brewer, Wesley Widener, Patrick Anantharaj, Valentine Wang, Feiyi Beck, Tom Shankar, Arjun Oral, Sarp |
| author_facet | Brewer, Wesley Widener, Patrick Anantharaj, Valentine Wang, Feiyi Beck, Tom Shankar, Arjun Oral, Sarp |
| contents | This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains - climate, nuclear fusion, bio/health, and materials - to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework composed of Data Readiness Levels (raw to AI-ready) and Data Processing Stages (ingest to shard), both tailored to high performance computing (HPC) environments. This framework outlines key challenges in transforming scientific data for scalable AI training, emphasizing transformer-based generative models. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_23018 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Data Readiness for Scientific AI at Scale Brewer, Wesley Widener, Patrick Anantharaj, Valentine Wang, Feiyi Beck, Tom Shankar, Arjun Oral, Sarp Artificial Intelligence Computational Engineering, Finance, and Science Distributed, Parallel, and Cluster Computing Machine Learning I.2.6 This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains - climate, nuclear fusion, bio/health, and materials - to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework composed of Data Readiness Levels (raw to AI-ready) and Data Processing Stages (ingest to shard), both tailored to high performance computing (HPC) environments. This framework outlines key challenges in transforming scientific data for scalable AI training, emphasizing transformer-based generative models. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science. |
| title | Data Readiness for Scientific AI at Scale |
| topic | Artificial Intelligence Computational Engineering, Finance, and Science Distributed, Parallel, and Cluster Computing Machine Learning I.2.6 |
| url | https://arxiv.org/abs/2507.23018 |