AirQualityBench: A Realistic Evaluation Benchmark for Global Air Quality Forecasting

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Xing, Wang, Xu, Zhang, Yudong, Zhao, Huilin, Zhou, Zhengyang, Wang, Yang
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915987407765504
author Xu, Xing
Wang, Xu
Zhang, Yudong
Zhao, Huilin
Zhou, Zhengyang
Wang, Yang
author_facet Xu, Xing
Wang, Xu
Zhang, Yudong
Zhao, Huilin
Zhou, Zhengyang
Wang, Yang
contents Air-quality forecasting models are commonly evaluated on regional, preprocessed, and normalized datasets, where missing observations are removed or artificially completed. Such protocols simplify comparison but hide the conditions that dominate real monitoring networks: uneven global coverage, structured missingness, heterogeneous pollutant scales, and deployment cost. We introduce \textbf{AirQualityBench}, a global multi-pollutant benchmark designed to evaluate forecasting models under these realistic conditions. The benchmark contains hourly observations from 3,720 monitoring stations over 2021--2025, covers six major pollutants, and preserves provider-native observation masks. Rather than imputing a dense data tensor, AirQualityBench exposes missingness as part of the forecasting problem and reports errors on valid future observations after inverse transformation to physical concentration scales. Evaluating representative spatio-temporal models under this unified protocol shows that strong performance on sanitized datasets does not reliably transfer to global, fragmented monitoring streams. AirQualityBench therefore serves as a realistic testbed for scalable, mask-aware, and physically interpretable air-quality forecasting. All benchmark data, code, evaluation scripts, and baseline implementations are available at \href{https://github.com/Star-Learning/AirQualityBench}{GitHub}.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05854
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AirQualityBench: A Realistic Evaluation Benchmark for Global Air Quality Forecasting
Xu, Xing
Wang, Xu
Zhang, Yudong
Zhao, Huilin
Zhou, Zhengyang
Wang, Yang
Artificial Intelligence
Air-quality forecasting models are commonly evaluated on regional, preprocessed, and normalized datasets, where missing observations are removed or artificially completed. Such protocols simplify comparison but hide the conditions that dominate real monitoring networks: uneven global coverage, structured missingness, heterogeneous pollutant scales, and deployment cost. We introduce \textbf{AirQualityBench}, a global multi-pollutant benchmark designed to evaluate forecasting models under these realistic conditions. The benchmark contains hourly observations from 3,720 monitoring stations over 2021--2025, covers six major pollutants, and preserves provider-native observation masks. Rather than imputing a dense data tensor, AirQualityBench exposes missingness as part of the forecasting problem and reports errors on valid future observations after inverse transformation to physical concentration scales. Evaluating representative spatio-temporal models under this unified protocol shows that strong performance on sanitized datasets does not reliably transfer to global, fragmented monitoring streams. AirQualityBench therefore serves as a realistic testbed for scalable, mask-aware, and physically interpretable air-quality forecasting. All benchmark data, code, evaluation scripts, and baseline implementations are available at \href{https://github.com/Star-Learning/AirQualityBench}{GitHub}.
title AirQualityBench: A Realistic Evaluation Benchmark for Global Air Quality Forecasting
topic Artificial Intelligence
url https://arxiv.org/abs/2605.05854