Metadata practices for simulation workflows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Villamar, José, Kelbling, Matthias, More, Heather L., Denker, Michael, Tetzlaff, Tom, Senk, Johanna, Thober, Stephan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913889711554560
author Villamar, José
Kelbling, Matthias
More, Heather L.
Denker, Michael
Tetzlaff, Tom
Senk, Johanna
Thober, Stephan
author_facet Villamar, José
Kelbling, Matthias
More, Heather L.
Denker, Michael
Tetzlaff, Tom
Senk, Johanna
Thober, Stephan
contents Computer simulations are an essential pillar of knowledge generation in science. Exploring, understanding, reproducing, and sharing the results of simulations relies on tracking and organizing the metadata describing the numerical experiments. The models used to understand real-world systems, and the computational machinery required to simulate them, are typically complex, and produce large amounts of heterogeneous metadata. Here, we present general practices for acquiring and handling metadata that are agnostic to software and hardware, and highly flexible for the user. These consist of two steps: 1) recording and storing raw metadata, and 2) selecting and structuring metadata. As a proof of concept, we develop the Archivist, a Python tool to help with the second step, and use it to apply our practices to distinct high-performance computing use cases from neuroscience and hydrology. Our practices and the Archivist can readily be applied to existing workflows without the need for substantial restructuring. They support sustainable numerical workflows, fostering replicability, reproducibility, data exploration, and data sharing in simulation-based research.
format Preprint
id arxiv_https___arxiv_org_abs_2408_17309
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Metadata practices for simulation workflows
Villamar, José
Kelbling, Matthias
More, Heather L.
Denker, Michael
Tetzlaff, Tom
Senk, Johanna
Thober, Stephan
Information Retrieval
Computer simulations are an essential pillar of knowledge generation in science. Exploring, understanding, reproducing, and sharing the results of simulations relies on tracking and organizing the metadata describing the numerical experiments. The models used to understand real-world systems, and the computational machinery required to simulate them, are typically complex, and produce large amounts of heterogeneous metadata. Here, we present general practices for acquiring and handling metadata that are agnostic to software and hardware, and highly flexible for the user. These consist of two steps: 1) recording and storing raw metadata, and 2) selecting and structuring metadata. As a proof of concept, we develop the Archivist, a Python tool to help with the second step, and use it to apply our practices to distinct high-performance computing use cases from neuroscience and hydrology. Our practices and the Archivist can readily be applied to existing workflows without the need for substantial restructuring. They support sustainable numerical workflows, fostering replicability, reproducibility, data exploration, and data sharing in simulation-based research.
title Metadata practices for simulation workflows
topic Information Retrieval
url https://arxiv.org/abs/2408.17309