Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yao, Wendong, Huang, Binhua, Dev, Soumyabrata
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918152155168768
author Yao, Wendong
Huang, Binhua
Dev, Soumyabrata
author_facet Yao, Wendong
Huang, Binhua
Dev, Soumyabrata
contents Forecasting high-resolution land subsidence is a critical yet challenging task due to its complex, non-linear dynamics. While standard architectures like ConvLSTM often fail to model long-range dependencies, we argue that a more fundamental limitation of prior work lies in the uni-modal data paradigm. To address this, we propose the Multi-Modal Spatio-Temporal Transformer (MM-STT), a novel framework that fuses dynamic displacement data with static physical priors. Its core innovation is a joint spatio-temporal attention mechanism that processes all multi-modal features in a unified manner. On the public EGMS dataset, MM-STT establishes a new state-of-the-art, reducing the long-range forecast RMSE by an order of magnitude compared to all baselines, including SOTA methods like STGCN and STAEformer. Our results demonstrate that for this class of problems, an architecture's inherent capacity for deep multi-modal fusion is paramount for achieving transformative performance.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25393
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction
Yao, Wendong
Huang, Binhua
Dev, Soumyabrata
Computer Vision and Pattern Recognition
Artificial Intelligence
Forecasting high-resolution land subsidence is a critical yet challenging task due to its complex, non-linear dynamics. While standard architectures like ConvLSTM often fail to model long-range dependencies, we argue that a more fundamental limitation of prior work lies in the uni-modal data paradigm. To address this, we propose the Multi-Modal Spatio-Temporal Transformer (MM-STT), a novel framework that fuses dynamic displacement data with static physical priors. Its core innovation is a joint spatio-temporal attention mechanism that processes all multi-modal features in a unified manner. On the public EGMS dataset, MM-STT establishes a new state-of-the-art, reducing the long-range forecast RMSE by an order of magnitude compared to all baselines, including SOTA methods like STGCN and STAEformer. Our results demonstrate that for this class of problems, an architecture's inherent capacity for deep multi-modal fusion is paramount for achieving transformative performance.
title Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.25393