Transformers vs. Recurrent Models for Estimating Forest Gross Primary Production

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Montero, David, Mahecha, Miguel D., Martinuzzi, Francesco, Aybar, César, Klosterhalfen, Anne, Knohl, Alexander, Anaya, Jesús, Mosig, Clemens, Wieneke, Sebastian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914159036203008
author Montero, David
Mahecha, Miguel D.
Martinuzzi, Francesco
Aybar, César
Klosterhalfen, Anne
Knohl, Alexander
Anaya, Jesús
Mosig, Clemens
Wieneke, Sebastian
author_facet Montero, David
Mahecha, Miguel D.
Martinuzzi, Francesco
Aybar, César
Klosterhalfen, Anne
Knohl, Alexander
Anaya, Jesús
Mosig, Clemens
Wieneke, Sebastian
contents Monitoring the spatiotemporal dynamics of forest CO$_2$ uptake (Gross Primary Production, GPP), remains a central challenge in terrestrial ecosystem research. While Eddy Covariance (EC) towers provide high-frequency estimates, their limited spatial coverage constrains large-scale assessments. Remote sensing offers a scalable alternative, yet most approaches rely on single-sensor spectral indices and statistical models that are often unable to capture the complex temporal dynamics of GPP. Recent advances in deep learning (DL) and data fusion offer new opportunities to better represent the temporal dynamics of vegetation processes, but comparative evaluations of state-of-the-art DL models for multimodal GPP prediction remain scarce. Here, we explore the performance of two representative models for predicting GPP: 1) GPT-2, a transformer architecture, and 2) Long Short-Term Memory (LSTM), a recurrent neural network, using multivariate inputs. Overall, both achieve similar accuracy. But, while LSTM performs better overall, GPT-2 excels during extreme events. Analysis of temporal context length further reveals that LSTM attains similar accuracy using substantially shorter input windows than GPT-2, highlighting an accuracy-efficiency trade-off between the two architectures. Feature importance analysis reveals radiation as the dominant predictor, followed by Sentinel-2, MODIS land surface temperature, and Sentinel-1 contributions. Our results demonstrate how model architecture, context length, and multimodal inputs jointly determine performance in GPP prediction, guiding future developments of DL frameworks for monitoring terrestrial carbon dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11880
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transformers vs. Recurrent Models for Estimating Forest Gross Primary Production
Montero, David
Mahecha, Miguel D.
Martinuzzi, Francesco
Aybar, César
Klosterhalfen, Anne
Knohl, Alexander
Anaya, Jesús
Mosig, Clemens
Wieneke, Sebastian
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Monitoring the spatiotemporal dynamics of forest CO$_2$ uptake (Gross Primary Production, GPP), remains a central challenge in terrestrial ecosystem research. While Eddy Covariance (EC) towers provide high-frequency estimates, their limited spatial coverage constrains large-scale assessments. Remote sensing offers a scalable alternative, yet most approaches rely on single-sensor spectral indices and statistical models that are often unable to capture the complex temporal dynamics of GPP. Recent advances in deep learning (DL) and data fusion offer new opportunities to better represent the temporal dynamics of vegetation processes, but comparative evaluations of state-of-the-art DL models for multimodal GPP prediction remain scarce. Here, we explore the performance of two representative models for predicting GPP: 1) GPT-2, a transformer architecture, and 2) Long Short-Term Memory (LSTM), a recurrent neural network, using multivariate inputs. Overall, both achieve similar accuracy. But, while LSTM performs better overall, GPT-2 excels during extreme events. Analysis of temporal context length further reveals that LSTM attains similar accuracy using substantially shorter input windows than GPT-2, highlighting an accuracy-efficiency trade-off between the two architectures. Feature importance analysis reveals radiation as the dominant predictor, followed by Sentinel-2, MODIS land surface temperature, and Sentinel-1 contributions. Our results demonstrate how model architecture, context length, and multimodal inputs jointly determine performance in GPP prediction, guiding future developments of DL frameworks for monitoring terrestrial carbon dynamics.
title Transformers vs. Recurrent Models for Estimating Forest Gross Primary Production
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.11880