Optimizing Cloud-to-GPU Throughput for Deep Learning With Earth Observation Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zaytar, Akram, Robinson, Caleb, Tadesse, Girmaw Abebe, Glazer, Tammy, Hacheme, Gilles, Ortiz, Anthony, Dodhia, Rahul M, Ferres, Juan M Lavista
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913881834651648
author Zaytar, Akram
Robinson, Caleb
Tadesse, Girmaw Abebe
Glazer, Tammy
Hacheme, Gilles
Ortiz, Anthony
Dodhia, Rahul M
Ferres, Juan M Lavista
author_facet Zaytar, Akram
Robinson, Caleb
Tadesse, Girmaw Abebe
Glazer, Tammy
Hacheme, Gilles
Ortiz, Anthony
Dodhia, Rahul M
Ferres, Juan M Lavista
contents Training deep learning models on petabyte-scale Earth observation (EO) data requires separating compute resources from data storage. However, standard PyTorch data loaders cannot keep modern GPUs utilized when streaming GeoTIFF files directly from cloud storage. In this work, we benchmark GeoTIFF loading throughput from both cloud object storage and local SSD, systematically testing different loader configurations and data parameters. We focus on tile-aligned reads and worker thread pools, using Bayesian optimization to find optimal settings for each storage type. Our optimized configurations increase remote data loading throughput by 20x and local throughput by 4x compared to default settings. On three public EO benchmarks, models trained with optimized remote loading achieve the same accuracy as local training within identical time budgets. We improve validation IoU by 6-15% and maintain 85-95% GPU utilization versus 0-30% with standard configurations. Code is publicly available at https://github.com/microsoft/pytorch-cloud-geotiff-optimization
format Preprint
id arxiv_https___arxiv_org_abs_2506_06235
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing Cloud-to-GPU Throughput for Deep Learning With Earth Observation Data
Zaytar, Akram
Robinson, Caleb
Tadesse, Girmaw Abebe
Glazer, Tammy
Hacheme, Gilles
Ortiz, Anthony
Dodhia, Rahul M
Ferres, Juan M Lavista
Computer Vision and Pattern Recognition
Training deep learning models on petabyte-scale Earth observation (EO) data requires separating compute resources from data storage. However, standard PyTorch data loaders cannot keep modern GPUs utilized when streaming GeoTIFF files directly from cloud storage. In this work, we benchmark GeoTIFF loading throughput from both cloud object storage and local SSD, systematically testing different loader configurations and data parameters. We focus on tile-aligned reads and worker thread pools, using Bayesian optimization to find optimal settings for each storage type. Our optimized configurations increase remote data loading throughput by 20x and local throughput by 4x compared to default settings. On three public EO benchmarks, models trained with optimized remote loading achieve the same accuracy as local training within identical time budgets. We improve validation IoU by 6-15% and maintain 85-95% GPU utilization versus 0-30% with standard configurations. Code is publicly available at https://github.com/microsoft/pytorch-cloud-geotiff-optimization
title Optimizing Cloud-to-GPU Throughput for Deep Learning With Earth Observation Data
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.06235