ShardTensor: Domain Parallelism for Scientific Machine Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Adams, Corey, Harrington, Peter, Subramaniam, Akshay, Abbas, Mohammad Shoaib, Pathak, Jaideep, Pritchard, Mike, Choudhry, Sanjay
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914555487059968
author Adams, Corey
Harrington, Peter
Subramaniam, Akshay
Abbas, Mohammad Shoaib
Pathak, Jaideep
Pritchard, Mike
Choudhry, Sanjay
author_facet Adams, Corey
Harrington, Peter
Subramaniam, Akshay
Abbas, Mohammad Shoaib
Pathak, Jaideep
Pritchard, Mike
Choudhry, Sanjay
contents Scientific Machine Learning (SciML) faces unique challenges for extreme-resolution data, with mitigations that often fail to scale or degrade the accuracy of trained models. While some specialized methods have achieved remarkable results in training models or performing inference on massive spatial datasets with bespoke techniques, there is no generalized framework for parallelization over input data below batch size one per device. In this work we introduce ShardTensor: a novel paradigm of domain parallelism that enables flexible scaling of input data to arbitrary sizes. By decoupling the spatial dimensionality of input data from hardware constraints, ShardTensor enables scientific machine learning workloads to reach new levels of high fidelity training and inference. We demonstrate both strong and weak scaling of workloads during training and inference, showing improved latency with strong scaling and demonstrating the capacity to process higher data sizes with weak scaling. Additionally, we demonstrate multiple dimensions of parallelization, removing barriers to SciML on extreme-scale inputs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11111
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ShardTensor: Domain Parallelism for Scientific Machine Learning
Adams, Corey
Harrington, Peter
Subramaniam, Akshay
Abbas, Mohammad Shoaib
Pathak, Jaideep
Pritchard, Mike
Choudhry, Sanjay
Distributed, Parallel, and Cluster Computing
Machine Learning
Scientific Machine Learning (SciML) faces unique challenges for extreme-resolution data, with mitigations that often fail to scale or degrade the accuracy of trained models. While some specialized methods have achieved remarkable results in training models or performing inference on massive spatial datasets with bespoke techniques, there is no generalized framework for parallelization over input data below batch size one per device. In this work we introduce ShardTensor: a novel paradigm of domain parallelism that enables flexible scaling of input data to arbitrary sizes. By decoupling the spatial dimensionality of input data from hardware constraints, ShardTensor enables scientific machine learning workloads to reach new levels of high fidelity training and inference. We demonstrate both strong and weak scaling of workloads during training and inference, showing improved latency with strong scaling and demonstrating the capacity to process higher data sizes with weak scaling. Additionally, we demonstrate multiple dimensions of parallelization, removing barriers to SciML on extreme-scale inputs.
title ShardTensor: Domain Parallelism for Scientific Machine Learning
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2605.11111