MultiMAE Meets Earth Observation: Pre-training Multi-modal Multi-task Masked Autoencoders for Earth Observation Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Sosa, Jose, Rukhovich, Danila, Kacem, Anis, Aouada, Djamila |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Effective is Pre-training of Large Masked Autoencoders for Downstream Earth Observation Tasks?
by: Sosa, Jose, et al.
Published: (2024)
by: Sosa, Jose, et al.
Published: (2024)
Enabling Training-Free Text-Based Remote Sensing Segmentation
by: Sosa, Jose, et al.
Published: (2026)
by: Sosa, Jose, et al.
Published: (2026)
CAD-Recode: Reverse Engineering CAD Code from Point Clouds
by: Rukhovich, Danila, et al.
Published: (2024)
by: Rukhovich, Danila, et al.
Published: (2024)
MiCADangelo: Fine-Grained Reconstruction of Constrained CAD Models from 3D Scans
by: Karadeniz, Ahmet Serdar, et al.
Published: (2025)
by: Karadeniz, Ahmet Serdar, et al.
Published: (2025)
MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition
by: Xiang, Peihao, et al.
Published: (2024)
by: Xiang, Peihao, et al.
Published: (2024)
Domain Adaptation for Multi-label Image Classification: a Discriminator-free Approach
by: Singh, Inder Pal, et al.
Published: (2025)
by: Singh, Inder Pal, et al.
Published: (2025)
MultiMAE for Brain MRIs: Robustness to Missing Inputs Using Multi-Modal Masked Autoencoder
by: Erdur, Ayhan Can, et al.
Published: (2025)
by: Erdur, Ayhan Can, et al.
Published: (2025)
CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task Solvers
by: Mallis, Dimitrios, et al.
Published: (2024)
by: Mallis, Dimitrios, et al.
Published: (2024)
Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detection
by: Nguyen, Dat, et al.
Published: (2025)
by: Nguyen, Dat, et al.
Published: (2025)
LAA-X: Unified Localized Artifact Attention for Quality-Agnostic and Generalizable Face Forgery Detection
by: Nguyen, Dat, et al.
Published: (2026)
by: Nguyen, Dat, et al.
Published: (2026)
NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining
by: Zeng, Liang, et al.
Published: (2026)
by: Zeng, Liang, et al.
Published: (2026)
Uncertainty-Aware Knowledge Distillation for Compact and Efficient 6DoF Pose Estimation
by: Ousalah, Nassim Ali, et al.
Published: (2025)
by: Ousalah, Nassim Ali, et al.
Published: (2025)
FPG-NAS: FLOPs-Aware Gated Differentiable Neural Architecture Search for Efficient 6DoF Pose Estimation
by: Ousalah, Nassim Ali, et al.
Published: (2025)
by: Ousalah, Nassim Ali, et al.
Published: (2025)
PICASSO: A Feed-Forward Framework for Parametric Inference of CAD Sketches via Rendering Self-Supervision
by: Karadeniz, Ahmet Serdar, et al.
Published: (2024)
by: Karadeniz, Ahmet Serdar, et al.
Published: (2024)
TerraMAE: Learning Spatial-Spectral Representations from Hyperspectral Earth Observation Data via Adaptive Masked Autoencoders
by: Faruk, Tanjim Bin, et al.
Published: (2025)
by: Faruk, Tanjim Bin, et al.
Published: (2025)
EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues
by: Soni, Sagar, et al.
Published: (2024)
by: Soni, Sagar, et al.
Published: (2024)
When Unsupervised Domain Adaptation meets One-class Anomaly Detection: Addressing the Two-fold Unsupervised Curse by Leveraging Anomaly Scarcity
by: Mejri, Nesryne, et al.
Published: (2025)
by: Mejri, Nesryne, et al.
Published: (2025)
TransCAD: A Hierarchical Transformer for CAD Sequence Inference from Point Clouds
by: Dupont, Elona, et al.
Published: (2024)
by: Dupont, Elona, et al.
Published: (2024)
R-MAE: Regions Meet Masked Autoencoders
by: Nguyen, Duy-Kien, et al.
Published: (2023)
by: Nguyen, Duy-Kien, et al.
Published: (2023)
CAD-SIGNet: CAD Language Inference from Point Clouds using Layer-wise Sketch Instance Guided Attention
by: Khan, Mohammad Sadil, et al.
Published: (2024)
by: Khan, Mohammad Sadil, et al.
Published: (2024)
Multi-modal Multi-task Pre-training for Improved Point Cloud Understanding
by: Liu, Liwen, et al.
Published: (2025)
by: Liu, Liwen, et al.
Published: (2025)
MV2MAE: Multi-View Video Masked Autoencoders
by: Shah, Ketul, et al.
Published: (2024)
by: Shah, Ketul, et al.
Published: (2024)
Lightweight Metadata-Aware Mixture-of-Experts Masked Autoencoder for Earth Observation
by: Albughdadi, Mohanad
Published: (2025)
by: Albughdadi, Mohanad
Published: (2025)
DAVINCI: A Single-Stage Architecture for Constrained CAD Sketch Inference
by: Karadeniz, Ahmet Serdar, et al.
Published: (2024)
by: Karadeniz, Ahmet Serdar, et al.
Published: (2024)
$\mathsf{CSMAE~}$:~Cataract Surgical Masked Autoencoder (MAE) based Pre-training
by: Shah, Nisarg A., et al.
Published: (2025)
by: Shah, Nisarg A., et al.
Published: (2025)
Cov2Pose: Leveraging Spatial Covariance for Direct Manifold-aware 6-DoF Object Pose Estimation
by: Ousalah, Nassim Ali, et al.
Published: (2026)
by: Ousalah, Nassim Ali, et al.
Published: (2026)
VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with Diffusion
by: Hicsonmez, Samet, et al.
Published: (2025)
by: Hicsonmez, Samet, et al.
Published: (2025)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025)
by: Mena, Francisco, et al.
Published: (2025)
Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
by: Mihai-Cristian, Pîrvu, et al.
Published: (2025)
by: Mihai-Cristian, Pîrvu, et al.
Published: (2025)
Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning
by: Ehsanpour, Mahsa, et al.
Published: (2024)
by: Ehsanpour, Mahsa, et al.
Published: (2024)
LAA-Net: Localized Artifact Attention Network for Quality-Agnostic and Generalizable Deepfake Detection
by: Nguyen, Dat, et al.
Published: (2024)
by: Nguyen, Dat, et al.
Published: (2024)
Text-to-CAD Evaluation with CADTests
by: Mallis, Dimitrios, et al.
Published: (2026)
by: Mallis, Dimitrios, et al.
Published: (2026)
Multi-label Image Classification using Adaptive Graph Convolutional Networks: from a Single Domain to Multiple Domains
by: Singh, Indel Pal, et al.
Published: (2023)
by: Singh, Indel Pal, et al.
Published: (2023)
EO-VAE: Towards A Multi-sensor Tokenizer for Earth Observation Data
by: Lehmann, Nils, et al.
Published: (2026)
by: Lehmann, Nils, et al.
Published: (2026)
Multi-Label Guided Soft Contrastive Learning for Efficient Earth Observation Pretraining
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
UniDet3D: Multi-dataset Indoor 3D Object Detection
by: Kolodiazhnyi, Maksim, et al.
Published: (2024)
by: Kolodiazhnyi, Maksim, et al.
Published: (2024)
EarthNets: Empowering AI in Earth Observation
by: Xiong, Zhitong, et al.
Published: (2022)
by: Xiong, Zhitong, et al.
Published: (2022)
cadrille: Multi-modal CAD Reconstruction with Reinforcement Learning
by: Kolodiazhnyi, Maksim, et al.
Published: (2025)
by: Kolodiazhnyi, Maksim, et al.
Published: (2025)
Hybrid Attention for Robust RGB-T Pedestrian Detection in Real-World Conditions
by: Rathinam, Arunkumar, et al.
Published: (2024)
by: Rathinam, Arunkumar, et al.
Published: (2024)
BigEarthNet.txt: A Large-Scale Multi-Sensor Image-Text Dataset and Benchmark for Earth Observation
by: Herzog, Johann-Ludwig, et al.
Published: (2026)
by: Herzog, Johann-Ludwig, et al.
Published: (2026)
Similar Items
-
How Effective is Pre-training of Large Masked Autoencoders for Downstream Earth Observation Tasks?
by: Sosa, Jose, et al.
Published: (2024) -
Enabling Training-Free Text-Based Remote Sensing Segmentation
by: Sosa, Jose, et al.
Published: (2026) -
CAD-Recode: Reverse Engineering CAD Code from Point Clouds
by: Rukhovich, Danila, et al.
Published: (2024) -
MiCADangelo: Fine-Grained Reconstruction of Constrained CAD Models from 3D Scans
by: Karadeniz, Ahmet Serdar, et al.
Published: (2025) -
MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition
by: Xiang, Peihao, et al.
Published: (2024)