MLCommons Cloud Masking Benchmark with Early Stopping

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chennamsetti, Varshitha, von Laszewski, Gregor, Gu, Ruochen, Mehnaz, Laiba, Papay, Juri, Jackson, Samuel, Thiyagalingam, Jeyan, Samsonau, Sergey V., Fox, Geoffrey C.
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916272146481152
author Chennamsetti, Varshitha
von Laszewski, Gregor
Gu, Ruochen
Mehnaz, Laiba
Papay, Juri
Jackson, Samuel
Thiyagalingam, Jeyan
Samsonau, Sergey V.
Fox, Geoffrey C.
author_facet Chennamsetti, Varshitha
von Laszewski, Gregor
Gu, Ruochen
Mehnaz, Laiba
Papay, Juri
Jackson, Samuel
Thiyagalingam, Jeyan
Samsonau, Sergey V.
Fox, Geoffrey C.
contents In this paper, we report on work performed for the MLCommons Science Working Group on the cloud masking benchmark. MLCommons is a consortium that develops and maintains several scientific benchmarks that aim to benefit developments in AI. The benchmarks are conducted on the High Performance Computing (HPC) Clusters of New York University and University of Virginia, as well as a commodity desktop. We provide a description of the cloud masking benchmark, as well as a summary of our submission to MLCommons on the benchmark experiment we conducted. It includes a modification to the reference implementation of the cloud masking benchmark enabling early stopping. This benchmark is executed on the NYU HPC through a custom batch script that runs the various experiments through the batch queuing system while allowing for variation on the number of epochs trained. Our submission includes the modified code, a custom batch script to modify epochs, documentation, and the benchmark results. We report the highest accuracy (scientific metric) and the average time taken (performance metric) for training and inference that was achieved on NYU HPC Greene. We also provide a comparison of the compute capabilities between different systems by running the benchmark for one epoch. Our submission can be found in a Globus repository that is accessible to MLCommons Science Working Group.
format Preprint
id arxiv_https___arxiv_org_abs_2401_08636
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MLCommons Cloud Masking Benchmark with Early Stopping
Chennamsetti, Varshitha
von Laszewski, Gregor
Gu, Ruochen
Mehnaz, Laiba
Papay, Juri
Jackson, Samuel
Thiyagalingam, Jeyan
Samsonau, Sergey V.
Fox, Geoffrey C.
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
In this paper, we report on work performed for the MLCommons Science Working Group on the cloud masking benchmark. MLCommons is a consortium that develops and maintains several scientific benchmarks that aim to benefit developments in AI. The benchmarks are conducted on the High Performance Computing (HPC) Clusters of New York University and University of Virginia, as well as a commodity desktop. We provide a description of the cloud masking benchmark, as well as a summary of our submission to MLCommons on the benchmark experiment we conducted. It includes a modification to the reference implementation of the cloud masking benchmark enabling early stopping. This benchmark is executed on the NYU HPC through a custom batch script that runs the various experiments through the batch queuing system while allowing for variation on the number of epochs trained. Our submission includes the modified code, a custom batch script to modify epochs, documentation, and the benchmark results. We report the highest accuracy (scientific metric) and the average time taken (performance metric) for training and inference that was achieved on NYU HPC Greene. We also provide a comparison of the compute capabilities between different systems by running the benchmark for one epoch. Our submission can be found in a Globus repository that is accessible to MLCommons Science Working Group.
title MLCommons Cloud Masking Benchmark with Early Stopping
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2401.08636