Distributed Cross-Channel Hierarchical Aggregation for Foundation Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tsaris, Aristeidis, Lyngaas, Isaac, Lagregren, John, Wahib, Mohamed, York, Larry, Balaprakash, Prasanna, Lu, Dan, Wang, Feiyi, Wang, Xiao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908423965114368
author Tsaris, Aristeidis
Lyngaas, Isaac
Lagregren, John
Wahib, Mohamed
York, Larry
Balaprakash, Prasanna
Lu, Dan
Wang, Feiyi
Wang, Xiao
author_facet Tsaris, Aristeidis
Lyngaas, Isaac
Lagregren, John
Wahib, Mohamed
York, Larry
Balaprakash, Prasanna
Lu, Dan
Wang, Feiyi
Wang, Xiao
contents Vision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from their ability to aggregate images from diverse sources such as varying physical groundings or data acquisition systems and to learn spatio-temporal correlations using transformer architectures. However, tokenizing and aggregating images can be compute-intensive, a challenge not fully addressed by current distributed methods. In this work, we introduce the Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) approach designed for datasets with a large number of channels across image modalities. Our method is compatible with any model-parallel strategy and any type of vision transformer architecture, significantly improving computational efficiency. We evaluated D-CHAG on hyperspectral imaging and weather forecasting tasks. When integrated with tensor parallelism and model sharding, our approach achieved up to a 75% reduction in memory usage and more than doubled sustained throughput on up to 1,024 AMD GPUs on the Frontier Supercomputer.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21411
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Distributed Cross-Channel Hierarchical Aggregation for Foundation Models
Tsaris, Aristeidis
Lyngaas, Isaac
Lagregren, John
Wahib, Mohamed
York, Larry
Balaprakash, Prasanna
Lu, Dan
Wang, Feiyi
Wang, Xiao
Machine Learning
Vision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from their ability to aggregate images from diverse sources such as varying physical groundings or data acquisition systems and to learn spatio-temporal correlations using transformer architectures. However, tokenizing and aggregating images can be compute-intensive, a challenge not fully addressed by current distributed methods. In this work, we introduce the Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) approach designed for datasets with a large number of channels across image modalities. Our method is compatible with any model-parallel strategy and any type of vision transformer architecture, significantly improving computational efficiency. We evaluated D-CHAG on hyperspectral imaging and weather forecasting tasks. When integrated with tensor parallelism and model sharding, our approach achieved up to a 75% reduction in memory usage and more than doubled sustained throughput on up to 1,024 AMD GPUs on the Frontier Supercomputer.
title Distributed Cross-Channel Hierarchical Aggregation for Foundation Models
topic Machine Learning
url https://arxiv.org/abs/2506.21411