Jigsaw: Training Multi-Billion-Parameter AI Weather Models with Optimized Model Parallelism
Fuente:
arXiv
Saved in:
| Main Authors: | Kieckhefen, Deifilia, Götz, Markus, Heyen, Lars H., Streit, Achim, Debus, Charlotte |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Energy Consumption in Parallel Neural Network Training
by: Huber, Philipp, et al.
Published: (2025)
by: Huber, Philipp, et al.
Published: (2025)
Sampling Parallelism for Fast and Efficient Bayesian Learning
by: Özdemir, Asena Karolin, et al.
Published: (2026)
by: Özdemir, Asena Karolin, et al.
Published: (2026)
Bayesian Lottery Ticket Hypothesis
by: Kuhn, Nicholas, et al.
Published: (2026)
by: Kuhn, Nicholas, et al.
Published: (2026)
Feed-Forward Optimization With Delayed Feedback for Neural Network Training
by: Flügel, Katharina, et al.
Published: (2023)
by: Flügel, Katharina, et al.
Published: (2023)
Beyond Backpropagation: Optimization with Multi-Tangent Forward Gradients
by: Flügel, Katharina, et al.
Published: (2024)
by: Flügel, Katharina, et al.
Published: (2024)
Model Fusion via Neuron Transplantation
by: Öz, Muhammed, et al.
Published: (2025)
by: Öz, Muhammed, et al.
Published: (2025)
Differentiable Power-Flow Optimization
by: Öz, Muhammed, et al.
Published: (2026)
by: Öz, Muhammed, et al.
Published: (2026)
Harnessing Orthogonality to Train Low-Rank Neural Networks
by: Coquelin, Daniel, et al.
Published: (2024)
by: Coquelin, Daniel, et al.
Published: (2024)
A Comparative Study of Pruning Methods in Transformer-based Time Series Forecasting
by: Kiefer, Nicholas, et al.
Published: (2024)
by: Kiefer, Nicholas, et al.
Published: (2024)
Inverse Design of Optical Multilayer Thin Films using Robust Masked Diffusion Models
by: Schaible, Jonas, et al.
Published: (2026)
by: Schaible, Jonas, et al.
Published: (2026)
AB-Training: A Communication-Efficient Approach for Distributed Low-Rank Learning
by: Coquelin, Daniel, et al.
Published: (2024)
by: Coquelin, Daniel, et al.
Published: (2024)
ReCycle: Fast and Efficient Long Time Series Forecasting with Residual Cyclic Transformers
by: Weyrauch, Arvid, et al.
Published: (2024)
by: Weyrauch, Arvid, et al.
Published: (2024)
PETNet -- Coincident Particle Event Detection using Spiking Neural Networks
by: Debus, Jan, et al.
Published: (2025)
by: Debus, Jan, et al.
Published: (2025)
Massively Parallel Genetic Optimization through Asynchronous Propagation of Populations
by: Taubert, Oskar, et al.
Published: (2023)
by: Taubert, Oskar, et al.
Published: (2023)
Sublinear Variational Optimization of Gaussian Mixture Models with Millions to Billions of Parameters
by: Salwig, Sebastian, et al.
Published: (2025)
by: Salwig, Sebastian, et al.
Published: (2025)
SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for Edge AI and Low-Resource Environments
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2024)
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2024)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
by: Hu, Zhengding, et al.
Published: (2026)
by: Hu, Zhengding, et al.
Published: (2026)
Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
by: Chaubard, Francois, et al.
Published: (2025)
by: Chaubard, Francois, et al.
Published: (2025)
Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials
by: Zhou, Yuanchang, et al.
Published: (2026)
by: Zhou, Yuanchang, et al.
Published: (2026)
Jigsaw Game: Federated Clustering
by: Xu, Jinxuan, et al.
Published: (2024)
by: Xu, Jinxuan, et al.
Published: (2024)
Jigsaw: Supporting Designers to Prototype Multimodal Applications by Chaining AI Foundation Models
by: Lin, David Chuan-En, et al.
Published: (2023)
by: Lin, David Chuan-En, et al.
Published: (2023)
Billion-Scale Graph Foundation Models
by: Bechler-Speicher, Maya, et al.
Published: (2026)
by: Bechler-Speicher, Maya, et al.
Published: (2026)
GroverGPT: A Large Language Model with 8 Billion Parameters for Quantum Searching
by: Wang, Haoran, et al.
Published: (2024)
by: Wang, Haoran, et al.
Published: (2024)
Global Vegetation Modeling with Pre-Trained Weather Transformers
by: Janetzky, Pascal, et al.
Published: (2024)
by: Janetzky, Pascal, et al.
Published: (2024)
A Scalable Multi-Task Model for Virtual Sensors
by: Götz, Leon, et al.
Published: (2026)
by: Götz, Leon, et al.
Published: (2026)
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
by: Ranjan, Aditya K., et al.
Published: (2025)
by: Ranjan, Aditya K., et al.
Published: (2025)
Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models
by: Cao, Shilei, et al.
Published: (2025)
by: Cao, Shilei, et al.
Published: (2025)
Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction
by: Hess, Florian, et al.
Published: (2026)
by: Hess, Florian, et al.
Published: (2026)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
by: Tuteja, Keshvi, et al.
Published: (2025)
by: Tuteja, Keshvi, et al.
Published: (2025)
Federated Full-Parameter Tuning of Billion-Sized Language Models with Communication Cost under 18 Kilobytes
by: Qin, Zhen, et al.
Published: (2023)
by: Qin, Zhen, et al.
Published: (2023)
Protocol Models: Scaling Decentralized Training with Communication-Efficient Model Parallelism
by: Ramasinghe, Sameera, et al.
Published: (2025)
by: Ramasinghe, Sameera, et al.
Published: (2025)
Hierarchical Sparse Circuit Extraction from Billion-Parameter Language Models through Scalable Attribution Graph Decomposition
by: Uddin, Mohammed Mudassir, et al.
Published: (2026)
by: Uddin, Mohammed Mudassir, et al.
Published: (2026)
TimeHF: Billion-Scale Time Series Models Guided by Human Feedback
by: Qi, Yongzhi, et al.
Published: (2025)
by: Qi, Yongzhi, et al.
Published: (2025)
Shaping AI's Impact on Billions of Lives
by: Cuéllar, Mariano-Florentino, et al.
Published: (2024)
by: Cuéllar, Mariano-Florentino, et al.
Published: (2024)
Wavelet Latent Diffusion (Wala): Billion-Parameter 3D Generative Model with Compact Wavelet Encodings
by: Sanghi, Aditya, et al.
Published: (2024)
by: Sanghi, Aditya, et al.
Published: (2024)
A Comparative Analysis of Adversarial Robustness for Quantum and Classical Machine Learning Models
by: Wendlinger, Maximilian, et al.
Published: (2024)
by: Wendlinger, Maximilian, et al.
Published: (2024)
Reasoning Under 1 Billion: Memory-Augmented Reinforcement Learning for Large Language Models
by: Le, Hung, et al.
Published: (2025)
by: Le, Hung, et al.
Published: (2025)
Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language Models
by: Pham, Cuong, et al.
Published: (2025)
by: Pham, Cuong, et al.
Published: (2025)
On Optimizing the Communication of Model Parallelism
by: Zhuang, Yonghao, et al.
Published: (2022)
by: Zhuang, Yonghao, et al.
Published: (2022)
Zenith: Scaling up Ranking Models for Billion-scale Livestreaming Recommendation
by: Zhang, Ruifeng, et al.
Published: (2026)
by: Zhang, Ruifeng, et al.
Published: (2026)
Similar Items
-
Energy Consumption in Parallel Neural Network Training
by: Huber, Philipp, et al.
Published: (2025) -
Sampling Parallelism for Fast and Efficient Bayesian Learning
by: Özdemir, Asena Karolin, et al.
Published: (2026) -
Bayesian Lottery Ticket Hypothesis
by: Kuhn, Nicholas, et al.
Published: (2026) -
Feed-Forward Optimization With Delayed Feedback for Neural Network Training
by: Flügel, Katharina, et al.
Published: (2023) -
Beyond Backpropagation: Optimization with Multi-Tangent Forward Gradients
by: Flügel, Katharina, et al.
Published: (2024)