Distributed Cross-Channel Hierarchical Aggregation for Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tsaris, Aristeidis, Lyngaas, Isaac, Lagregren, John, Wahib, Mohamed, York, Larry, Balaprakash, Prasanna, Lu, Dan, Wang, Feiyi, Wang, Xiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sequence Length Scaling in Vision Transformers for Scientific Images on Frontier
by: Tsaris, Aristeidis, et al.
Published: (2024)
by: Tsaris, Aristeidis, et al.
Published: (2024)
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
FOCUS: Fused Observation of Channels for Unveiling Spectra
by: Xiao, Xi, et al.
Published: (2025)
by: Xiao, Xi, et al.
Published: (2025)
Pretraining Billion-scale Geospatial Foundational Models on Frontier
by: Tsaris, Aristeidis, et al.
Published: (2024)
by: Tsaris, Aristeidis, et al.
Published: (2024)
ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Scalable Artificial Intelligence for Science: Perspectives, Methods and Exemplars
by: Brewer, Wesley, et al.
Published: (2024)
by: Brewer, Wesley, et al.
Published: (2024)
Adaptive Patching for High-resolution Image Segmentation with Transformers
by: Zhang, Enzhi, et al.
Published: (2024)
by: Zhang, Enzhi, et al.
Published: (2024)
Towards Scaling Law Analysis For Spatiotemporal Weather Data
by: Kiefer, Alexander, et al.
Published: (2026)
by: Kiefer, Alexander, et al.
Published: (2026)
Large Language Models Inference Engines based on Spiking Neural Networks
by: Balaji, Adarsha, et al.
Published: (2025)
by: Balaji, Adarsha, et al.
Published: (2025)
Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training
by: Tyagi, Sahil, et al.
Published: (2026)
by: Tyagi, Sahil, et al.
Published: (2026)
The Unreasonable Effectiveness Of Early Discarding After One Epoch In Neural Network Hyperparameter Optimization
by: Egele, Romain, et al.
Published: (2024)
by: Egele, Romain, et al.
Published: (2024)
Bayesian optimized deep ensemble for uncertainty quantification of deep neural networks: a system safety case study on sodium fast reactor thermal stratification modeling
by: Abulawi, Zaid, et al.
Published: (2024)
by: Abulawi, Zaid, et al.
Published: (2024)
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
by: Dash, Sajal, et al.
Published: (2026)
by: Dash, Sajal, et al.
Published: (2026)
Network architecture search of X-ray based scientific applications
by: Balaji, Adarsha, et al.
Published: (2024)
by: Balaji, Adarsha, et al.
Published: (2024)
Decomposable Transformer Point Processes
by: Panos, Aristeidis
Published: (2024)
by: Panos, Aristeidis
Published: (2024)
The (R)evolution of Scientific Workflows in the Agentic AI Era: Towards Autonomous Science
by: Shin, Woong, et al.
Published: (2025)
by: Shin, Woong, et al.
Published: (2025)
Streamlining Ocean Dynamics Modeling with Fourier Neural Operators: A Multiobjective Hyperparameter and Architecture Optimization Approach
by: Sun, Yixuan, et al.
Published: (2024)
by: Sun, Yixuan, et al.
Published: (2024)
Uncertainty Quantification for Molecular Property Predictions with Graph Neural Architecture Search
by: Jiang, Shengli, et al.
Published: (2023)
by: Jiang, Shengli, et al.
Published: (2023)
Scaling Laws of Graph Neural Networks for Atomistic Materials Modeling
by: Li, Chaojian, et al.
Published: (2025)
by: Li, Chaojian, et al.
Published: (2025)
Attacking Attention of Foundation Models Disrupts Downstream Tasks
by: Silva, Hondamunige Prasanna, et al.
Published: (2025)
by: Silva, Hondamunige Prasanna, et al.
Published: (2025)
Global Attention with Linear Complexity for Exascale Generative Data Assimilation in Earth System Prediction
by: Wang, Xiao, et al.
Published: (2026)
by: Wang, Xiao, et al.
Published: (2026)
Out-of-Distribution Generalization in Graph Foundation Models
by: Li, Haoyang, et al.
Published: (2026)
by: Li, Haoyang, et al.
Published: (2026)
Bi-level Personalization for Federated Foundation Models: A Task-vector Aggregation Approach
by: Yang, Yiyuan, et al.
Published: (2025)
by: Yang, Yiyuan, et al.
Published: (2025)
Coding-Enforced Resilient and Secure Aggregation for Hierarchical Federated Learning
by: Weng, Shudi, et al.
Published: (2026)
by: Weng, Shudi, et al.
Published: (2026)
Generalizable Prediction Model of Molten Salt Mixture Density with Chemistry-Informed Transfer Learning
by: Barra, Julian, et al.
Published: (2024)
by: Barra, Julian, et al.
Published: (2024)
Improving Normative Modeling for Multi-modal Neuroimaging Data using mixture-of-product-of-experts variational autoencoders
by: Kumar, Sayantan, et al.
Published: (2023)
by: Kumar, Sayantan, et al.
Published: (2023)
SpikeWFM: Spiking-Aided Wireless Foundation Model for Robust Channel Prediction
by: Jing, Liwen, et al.
Published: (2026)
by: Jing, Liwen, et al.
Published: (2026)
Sequential Bayesian Neural Subnetwork Ensembles
by: Jantre, Sanket, et al.
Published: (2022)
by: Jantre, Sanket, et al.
Published: (2022)
Physics-Informed Heterogeneous Graph Neural Networks for DC Blocker Placement
by: Jin, Hongwei, et al.
Published: (2024)
by: Jin, Hongwei, et al.
Published: (2024)
dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence Learning
by: Shah, Arnav, et al.
Published: (2026)
by: Shah, Arnav, et al.
Published: (2026)
Single-Channel Tissue Segmentation via Cross-Modal Distillation from Foundation Models
by: Mohammad, Sakib, et al.
Published: (2026)
by: Mohammad, Sakib, et al.
Published: (2026)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025)
by: Tang, Shixiang, et al.
Published: (2025)
Federated Adapter on Foundation Models: An Out-Of-Distribution Approach
by: Yang, Yiyuan, et al.
Published: (2025)
by: Yang, Yiyuan, et al.
Published: (2025)
Tadashi: Enabling AI-Based Automated Code Generation With Guaranteed Correctness
by: Vatai, Emil, et al.
Published: (2024)
by: Vatai, Emil, et al.
Published: (2024)
Subgraph Aggregation for Out-of-Distribution Generalization on Graphs
by: Liu, Bowen, et al.
Published: (2024)
by: Liu, Bowen, et al.
Published: (2024)
FedSpy-LLM: Towards Scalable and Generalizable Data Reconstruction Attacks from Gradients on LLMs
by: Meerza, Syed Irfan Ali, et al.
Published: (2026)
by: Meerza, Syed Irfan Ali, et al.
Published: (2026)
How does ion temperature gradient turbulence depend on magnetic geometry? Insights from data and machine learning
by: Landreman, Matt, et al.
Published: (2025)
by: Landreman, Matt, et al.
Published: (2025)
Mixture of Thoughts: Learning to Aggregate What Experts Think, Not Just What They Say
by: Fein-Ashley, Jacob, et al.
Published: (2025)
by: Fein-Ashley, Jacob, et al.
Published: (2025)
Matrix-free Neural Preconditioner for the Dirac Operator in Lattice Gauge Theory
by: Sun, Yixuan, et al.
Published: (2025)
by: Sun, Yixuan, et al.
Published: (2025)
Towards Foundation Models for Consensus Rank Aggregation
by: Jin, Yijun, et al.
Published: (2026)
by: Jin, Yijun, et al.
Published: (2026)
Similar Items
-
Sequence Length Scaling in Vision Transformers for Scientific Images on Frontier
by: Tsaris, Aristeidis, et al.
Published: (2024) -
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling
by: Wang, Xiao, et al.
Published: (2025) -
FOCUS: Fused Observation of Channels for Unveiling Spectra
by: Xiao, Xi, et al.
Published: (2025) -
Pretraining Billion-scale Geospatial Foundational Models on Frontier
by: Tsaris, Aristeidis, et al.
Published: (2024) -
ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability
by: Wang, Xiao, et al.
Published: (2024)