Scaling Laws of Global Weather Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Yuejiang, Huang, Langwen, Calotoiu, Alexandru, Hoefler, Torsten
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917296668147712
author Yu, Yuejiang
Huang, Langwen
Calotoiu, Alexandru
Hoefler, Torsten
author_facet Yu, Yuejiang
Huang, Langwen
Calotoiu, Alexandru
Hoefler, Torsten
contents Data-driven models are revolutionizing weather forecasting. To optimize training efficiency and model performance, this paper analyzes empirical scaling laws within this domain. We investigate the relationship between model performance (validation loss) and three key factors: model size ($N$), dataset size ($D$), and compute budget ($C$). Across a range of models, we find that Aurora exhibits the strongest data-scaling behavior: increasing the training dataset by 10x reduces validation loss by up to 3.2x. GraphCast demonstrates the highest parameter efficiency, yet suffers from limited hardware utilization. Our compute-optimal analysis indicates that, under fixed compute budgets, allocating resources to longer training durations yields greater performance gains than increasing model size. Furthermore, we analyze model shape and uncover scaling behaviors that differ fundamentally from those observed in language models: weather forecasting models consistently favor increased width over depth. These findings suggest that future weather models should prioritize wider architectures and larger effective training datasets to maximize predictive performance.
format Preprint
id arxiv_https___arxiv_org_abs_2602_22962
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scaling Laws of Global Weather Models
Yu, Yuejiang
Huang, Langwen
Calotoiu, Alexandru
Hoefler, Torsten
Machine Learning
Data-driven models are revolutionizing weather forecasting. To optimize training efficiency and model performance, this paper analyzes empirical scaling laws within this domain. We investigate the relationship between model performance (validation loss) and three key factors: model size ($N$), dataset size ($D$), and compute budget ($C$). Across a range of models, we find that Aurora exhibits the strongest data-scaling behavior: increasing the training dataset by 10x reduces validation loss by up to 3.2x. GraphCast demonstrates the highest parameter efficiency, yet suffers from limited hardware utilization. Our compute-optimal analysis indicates that, under fixed compute budgets, allocating resources to longer training durations yields greater performance gains than increasing model size. Furthermore, we analyze model shape and uncover scaling behaviors that differ fundamentally from those observed in language models: weather forecasting models consistently favor increased width over depth. These findings suggest that future weather models should prioritize wider architectures and larger effective training datasets to maximize predictive performance.
title Scaling Laws of Global Weather Models
topic Machine Learning
url https://arxiv.org/abs/2602.22962