Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Wendong, Huang, Binhua, Dev, Soumyabrata |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A multimodal Transformer for InSAR-based ground deformation forecasting with cross-site generalization across Europe
by: Yao, Wendong, et al.
Published: (2025)
by: Yao, Wendong, et al.
Published: (2025)
A Deep Learning Approach for Spatio-Temporal Forecasting of InSAR Ground Deformation in Eastern Ireland
by: Yao, Wendong, et al.
Published: (2025)
by: Yao, Wendong, et al.
Published: (2025)
MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
by: Huang, Binhua, et al.
Published: (2025)
by: Huang, Binhua, et al.
Published: (2025)
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
by: Huang, Binhua, et al.
Published: (2025)
by: Huang, Binhua, et al.
Published: (2025)
MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
by: Huang, Binhua, et al.
Published: (2025)
by: Huang, Binhua, et al.
Published: (2025)
TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers
by: Wang, Guoxin, et al.
Published: (2025)
by: Wang, Guoxin, et al.
Published: (2025)
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification
by: Wang, Xiao, et al.
Published: (2026)
by: Wang, Xiao, et al.
Published: (2026)
DIFFUMA: High-Fidelity Spatio-Temporal Video Prediction via Dual-Path Mamba and Diffusion Enhancement
by: Xie, Xinyu, et al.
Published: (2025)
by: Xie, Xinyu, et al.
Published: (2025)
Attention-based Multi-modal Deep Learning Model of Spatio-temporal Crop Yield Prediction with Satellite, Soil and Climate Data
by: Shyam, Gopal Krishna, et al.
Published: (2026)
by: Shyam, Gopal Krishna, et al.
Published: (2026)
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2025)
by: Gu, Xin, et al.
Published: (2025)
Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
by: Liu, Tengfei, et al.
Published: (2024)
by: Liu, Tengfei, et al.
Published: (2024)
GATS: Gaussian Aware Temporal Scaling Transformer for Invariant 4D Spatio-Temporal Point Cloud Representation
by: Tian, Jiayi, et al.
Published: (2026)
by: Tian, Jiayi, et al.
Published: (2026)
Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents
by: Xu, Zhou, et al.
Published: (2026)
by: Xu, Zhou, et al.
Published: (2026)
Deformable Dynamic Convolution for Accurate yet Efficient Spatio-Temporal Traffic Prediction
by: Jin, Hyeonseok, et al.
Published: (2025)
by: Jin, Hyeonseok, et al.
Published: (2025)
Multi-scale Temporal Fusion Transformer for Incomplete Vehicle Trajectory Prediction
by: Liu, Zhanwen, et al.
Published: (2024)
by: Liu, Zhanwen, et al.
Published: (2024)
Human-Centric Video Anomaly Detection Through Spatio-Temporal Pose Tokenization and Transformer
by: Noghre, Ghazal Alinezhad, et al.
Published: (2024)
by: Noghre, Ghazal Alinezhad, et al.
Published: (2024)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
by: Huang, Wei-Jhe, et al.
Published: (2024)
by: Huang, Wei-Jhe, et al.
Published: (2024)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Multi-modal Deepfake Detection and Localization with FPN-Transformer
by: Zheng, Chende, et al.
Published: (2025)
by: Zheng, Chende, et al.
Published: (2025)
Multi-Scale Spatio-Temporal Graph Convolutional Network for Facial Expression Spotting
by: Deng, Yicheng, et al.
Published: (2024)
by: Deng, Yicheng, et al.
Published: (2024)
Mechanistic Learning with Guided Diffusion Models to Predict Spatio-Temporal Brain Tumor Growth
by: Laslo, Daria, et al.
Published: (2025)
by: Laslo, Daria, et al.
Published: (2025)
STAA: Spatio-Temporal Attention Attribution for Real-Time Interpreting Transformer-based Video Models
by: Wang, Zerui, et al.
Published: (2024)
by: Wang, Zerui, et al.
Published: (2024)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
by: Hyun, Jeongseok, et al.
Published: (2025)
by: Hyun, Jeongseok, et al.
Published: (2025)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
Vision-Core Guided Contrastive Learning for Balanced Multi-modal Prognosis Prediction of Stroke
by: Chen, Liren, et al.
Published: (2026)
by: Chen, Liren, et al.
Published: (2026)
BD-SAT: High-resolution Land Use Land Cover Dataset & Benchmark Results for Developing Division: Dhaka, BD
by: Paul, Ovi, et al.
Published: (2024)
by: Paul, Ovi, et al.
Published: (2024)
Reprogramming Vision Foundation Models for Spatio-Temporal Forecasting
by: Chen, Changlu, et al.
Published: (2025)
by: Chen, Changlu, et al.
Published: (2025)
Agentic Spatio-Temporal Grounding via Collaborative Reasoning
by: Zhao, Heng, et al.
Published: (2026)
by: Zhao, Heng, et al.
Published: (2026)
Uncertainty-Aware Test-Time Adaptation for Cross-Region Spatio-Temporal Fusion of Land Surface Temperature
by: Bouaziz, Sofiane, et al.
Published: (2026)
by: Bouaziz, Sofiane, et al.
Published: (2026)
Temporal Reversal Regularization for Spiking Neural Networks: Hybrid Spatio-Temporal Invariance for Generalization
by: Zuo, Lin, et al.
Published: (2024)
by: Zuo, Lin, et al.
Published: (2024)
TrajFlow: Multi-modal Motion Prediction via Flow Matching
by: Yan, Qi, et al.
Published: (2025)
by: Yan, Qi, et al.
Published: (2025)
Can Multi-modal (reasoning) LLMs work as deepfake detectors?
by: Ren, Simiao, et al.
Published: (2025)
by: Ren, Simiao, et al.
Published: (2025)
A Memory-Efficient Framework for Deformable Transformer with Neural Architecture Search
by: Mao, Wendong, et al.
Published: (2025)
by: Mao, Wendong, et al.
Published: (2025)
MetaOcc: Spatio-Temporal Fusion of Surround-View 4D Radar and Camera for 3D Occupancy Prediction with Dual Training Strategies
by: Yang, Long, et al.
Published: (2025)
by: Yang, Long, et al.
Published: (2025)
Rethinking Patient Education as Multi-turn Multi-modal Interaction
by: Yao, Zonghai, et al.
Published: (2026)
by: Yao, Zonghai, et al.
Published: (2026)
DMTrack: Spatio-Temporal Multimodal Tracking via Dual-Adapter
by: Li, Weihong, et al.
Published: (2025)
by: Li, Weihong, et al.
Published: (2025)
SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action Recognition
by: Huang, Wenbo, et al.
Published: (2024)
by: Huang, Wenbo, et al.
Published: (2024)
UnityGraph: Unified Learning of Spatio-temporal features for Multi-person Motion Prediction
by: Qu, Kehua, et al.
Published: (2024)
by: Qu, Kehua, et al.
Published: (2024)
Advancing Stroke Risk Prediction Using a Multi-modal Foundation Model
by: Delgrange, Camille, et al.
Published: (2024)
by: Delgrange, Camille, et al.
Published: (2024)
MulCPred: Learning Multi-modal Concepts for Explainable Pedestrian Action Prediction
by: Feng, Yan, et al.
Published: (2024)
by: Feng, Yan, et al.
Published: (2024)
Similar Items
-
A multimodal Transformer for InSAR-based ground deformation forecasting with cross-site generalization across Europe
by: Yao, Wendong, et al.
Published: (2025) -
A Deep Learning Approach for Spatio-Temporal Forecasting of InSAR Ground Deformation in Eastern Ireland
by: Yao, Wendong, et al.
Published: (2025) -
MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
by: Huang, Binhua, et al.
Published: (2025) -
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
by: Huang, Binhua, et al.
Published: (2025) -
MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
by: Huang, Binhua, et al.
Published: (2025)