Dataset and Model Code for SWAT+ and Machine Learning Assessment of Satellite Precipitation in Data-Scarce Regions

Fuente: Zenodo
Guardado en:
Detalles Bibliográficos
Autor principal: Almeida, Manuel
Formato: Recurso digital
Publicado: Zenodo 2025
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866902289929732096
author Almeida, Manuel
author_facet Almeida, Manuel
contents <p><strong>Description:</strong></p> <p>This repository contains the code and data for a comparative evaluation of physically based (SWAT+) and machine learning (ML) hydrological models applied to two contrasting watersheds in Angola. The study investigates model performance using three satellite and reanalysis precipitation datasets: MSWEP, IMERG, and ERA5-Land. Model evaluation was performed using two complementary approaches: traditional calibration–validation metrics and flow-duration-curve (FDC) statistics to capture minimum, mean, and maximum flow behavior.</p> <p>The SWAT+ models were executed with <strong>SWAT+ version 61.0.2</strong> (swatplus-61.0.2-ifx-win_amd64-Rel.exe) and calibrated using the SWAT+ Toolbox (v2.4.1) via an automated Latin Hypercube Sampling (LHS) scheme targeting the Nash–Sutcliffe Efficiency (NSE) objective function. Machine learning models—including Random Forest and Artificial Neural Networks—were trained with input features selected through permutation importance analysis and further refined by ANN optimization. Input features comprised lagged precipitation, relative humidity, air temperature, and cyclical transformations of month to account for seasonality.</p> <p>The datasets were split into training (51%), validation (13%), and testing (36%) subsets, ensuring robust evaluation across wet and dry periods. Data imbalance was mitigated using SMOGN resampling to improve predictions of rare and extreme flow events. Hyperparameter tuning was conducted via Hyperopt with up to 100 evaluations.</p> <p>Key findings highlight the variability in model skill depending on watershed size, flow regime, and evaluation method. MSWEP emerged as the most reliable precipitation dataset, while ML models better captured flow extremes, especially when enhanced by SMOGN. This work demonstrates the value of integrating physically based and data-driven models alongside multiple evaluation techniques to improve hydrological representation in data-scarce regions.</p> <p><strong>Included Files:</strong></p> <ul> <li> <p>SWAT+ model input files and calibration scripts</p> </li> <li> <p>Machine learning model scripts with feature selection, training, and evaluation code</p> </li> <li> <p>Preprocessing and SMOGN resampling implementation</p> </li> <li> <p>Hyperparameter optimization workflows using Hyperopt</p> </li> </ul> <p><strong>Software Requirements:</strong></p> <ul> <li> <p>SWAT+ version 61.0.2 (swatplus-61.0.2-ifx-win_amd64-Rel.exe)</p> </li> <li> <p>SWAT+ Toolbox (v2.4.1)</p> </li> <li> <p>Python (with libraries: scikit-learn, hyperopt, pandas, numpy, SMOGN implementation)</p> </li> </ul>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17736146
institution Zenodo
language
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Dataset and Model Code for SWAT+ and Machine Learning Assessment of Satellite Precipitation in Data-Scarce Regions
Almeida, Manuel
<p><strong>Description:</strong></p> <p>This repository contains the code and data for a comparative evaluation of physically based (SWAT+) and machine learning (ML) hydrological models applied to two contrasting watersheds in Angola. The study investigates model performance using three satellite and reanalysis precipitation datasets: MSWEP, IMERG, and ERA5-Land. Model evaluation was performed using two complementary approaches: traditional calibration–validation metrics and flow-duration-curve (FDC) statistics to capture minimum, mean, and maximum flow behavior.</p> <p>The SWAT+ models were executed with <strong>SWAT+ version 61.0.2</strong> (swatplus-61.0.2-ifx-win_amd64-Rel.exe) and calibrated using the SWAT+ Toolbox (v2.4.1) via an automated Latin Hypercube Sampling (LHS) scheme targeting the Nash–Sutcliffe Efficiency (NSE) objective function. Machine learning models—including Random Forest and Artificial Neural Networks—were trained with input features selected through permutation importance analysis and further refined by ANN optimization. Input features comprised lagged precipitation, relative humidity, air temperature, and cyclical transformations of month to account for seasonality.</p> <p>The datasets were split into training (51%), validation (13%), and testing (36%) subsets, ensuring robust evaluation across wet and dry periods. Data imbalance was mitigated using SMOGN resampling to improve predictions of rare and extreme flow events. Hyperparameter tuning was conducted via Hyperopt with up to 100 evaluations.</p> <p>Key findings highlight the variability in model skill depending on watershed size, flow regime, and evaluation method. MSWEP emerged as the most reliable precipitation dataset, while ML models better captured flow extremes, especially when enhanced by SMOGN. This work demonstrates the value of integrating physically based and data-driven models alongside multiple evaluation techniques to improve hydrological representation in data-scarce regions.</p> <p><strong>Included Files:</strong></p> <ul> <li> <p>SWAT+ model input files and calibration scripts</p> </li> <li> <p>Machine learning model scripts with feature selection, training, and evaluation code</p> </li> <li> <p>Preprocessing and SMOGN resampling implementation</p> </li> <li> <p>Hyperparameter optimization workflows using Hyperopt</p> </li> </ul> <p><strong>Software Requirements:</strong></p> <ul> <li> <p>SWAT+ version 61.0.2 (swatplus-61.0.2-ifx-win_amd64-Rel.exe)</p> </li> <li> <p>SWAT+ Toolbox (v2.4.1)</p> </li> <li> <p>Python (with libraries: scikit-learn, hyperopt, pandas, numpy, SMOGN implementation)</p> </li> </ul>
title Dataset and Model Code for SWAT+ and Machine Learning Assessment of Satellite Precipitation in Data-Scarce Regions
url https://doi.org/10.5281/zenodo.17736146