MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Nedungadi, Vishal, Kariryaa, Ankit, Oehmcke, Stefan, Belongie, Serge, Igel, Christian, Lang, Nico |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
by: Gordon, Lucia, et al.
Published: (2026)
by: Gordon, Lucia, et al.
Published: (2026)
Nacala-Roof-Material: Drone Imagery for Roof Detection, Classification, and Segmentation to Support Mosquito-borne Disease Risk Assessment
by: Guthula, Venkanna Babu, et al.
Published: (2024)
by: Guthula, Venkanna Babu, et al.
Published: (2024)
Familiarity-Based Open-Set Recognition Under Adversarial Attacks
by: Enevoldsen, Philip, et al.
Published: (2023)
by: Enevoldsen, Philip, et al.
Published: (2023)
SuperF: Neural Implicit Fields for Multi-Image Super-Resolution
by: Jyhne, Sander Riisøen, et al.
Published: (2025)
by: Jyhne, Sander Riisøen, et al.
Published: (2025)
Labeled Data Selection for Category Discovery
by: Zhao, Bingchen, et al.
Published: (2024)
by: Zhao, Bingchen, et al.
Published: (2024)
SSL-Interactions: Pretext Tasks for Interactive Trajectory Prediction
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
Multimodal classification of forest biodiversity potential from 2D orthophotos and 3D airborne laser scanning point clouds
by: Jensen, Simon B., et al.
Published: (2025)
by: Jensen, Simon B., et al.
Published: (2025)
Progressive Pretext Task Learning for Human Trajectory Prediction
by: Lin, Xiaotong, et al.
Published: (2024)
by: Lin, Xiaotong, et al.
Published: (2024)
Unlearning-based Neural Interpretations
by: Choi, Ching Lam, et al.
Published: (2024)
by: Choi, Ching Lam, et al.
Published: (2024)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
by: Pach, Mateusz, et al.
Published: (2025)
by: Pach, Mateusz, et al.
Published: (2025)
Taxonomy-Aware Evaluation of Vision-Language Models
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025)
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025)
Learning to Perceive "Where": Spatial Pretext Tasks for Robust Self-Supervised Learning
by: Shen, Yang, et al.
Published: (2026)
by: Shen, Yang, et al.
Published: (2026)
Affinity-Graph-Guided Contractive Learning for Pretext-Free Medical Image Segmentation with Minimal Annotation
by: Cheng, Zehua, et al.
Published: (2024)
by: Cheng, Zehua, et al.
Published: (2024)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
by: Yu, Hong-Tao, et al.
Published: (2025)
by: Yu, Hong-Tao, et al.
Published: (2025)
Better Language Models Exhibit Higher Visual Alignment
by: Ruthardt, Jona, et al.
Published: (2024)
by: Ruthardt, Jona, et al.
Published: (2024)
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
by: Pach, Mateusz, et al.
Published: (2026)
by: Pach, Mateusz, et al.
Published: (2026)
Compressing Vision Transformers in Geospatial Transfer Learning with Manifold-Constrained Optimization
by: Snyder, Thomas, et al.
Published: (2026)
by: Snyder, Thomas, et al.
Published: (2026)
Discriminative Class Tokens for Text-to-Image Diffusion Models
by: Schwartz, Idan, et al.
Published: (2023)
by: Schwartz, Idan, et al.
Published: (2023)
Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
Tree Counting by Bridging 3D Point Clouds with Imagery
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
Geospatial Representation Learning: A Survey from Deep Learning to The LLM Era
by: Hao, Xixuan, et al.
Published: (2025)
by: Hao, Xixuan, et al.
Published: (2025)
Time2Agri: Temporal Pretext Tasks for Agricultural Monitoring
by: Gupta, Moti Rattan, et al.
Published: (2025)
by: Gupta, Moti Rattan, et al.
Published: (2025)
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
by: Roberts, Jonathan, et al.
Published: (2023)
by: Roberts, Jonathan, et al.
Published: (2023)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
by: Wang, Tianyi, et al.
Published: (2025)
by: Wang, Tianyi, et al.
Published: (2025)
Coarse-To-Fine Tensor Trains for Compact Visual Representations
by: Loeschcke, Sebastian, et al.
Published: (2024)
by: Loeschcke, Sebastian, et al.
Published: (2024)
Geo2Vec: Shape- and Distance-Aware Neural Representation of Geospatial Entities
by: Chu, Chen, et al.
Published: (2025)
by: Chu, Chen, et al.
Published: (2025)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
by: Kim, Jeonghyeon, et al.
Published: (2025)
by: Kim, Jeonghyeon, et al.
Published: (2025)
From Pretext to Purpose: Batch-Adaptive Self-Supervised Learning
by: Zhang, Jiansong, et al.
Published: (2023)
by: Zhang, Jiansong, et al.
Published: (2023)
AltChart: Enhancing VLM-based Chart Summarization Through Multi-Pretext Tasks
by: Moured, Omar, et al.
Published: (2024)
by: Moured, Omar, et al.
Published: (2024)
Behavior-Grounded Lane Representation Learning for Multi-Task Traffic Digital Twins
by: Tamaru, Rei, et al.
Published: (2026)
by: Tamaru, Rei, et al.
Published: (2026)
PhysConvex: Physics-Informed 3D Dynamic Convex Radiance Fields for Reconstruction and Simulation
by: Wang, Dan, et al.
Published: (2026)
by: Wang, Dan, et al.
Published: (2026)
HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement
by: Schouten, Marco, et al.
Published: (2026)
by: Schouten, Marco, et al.
Published: (2026)
Learning Content-Aware Multi-Modal Joint Input Pruning via Bird's-Eye-View Representation
by: Li, Yuxin, et al.
Published: (2024)
by: Li, Yuxin, et al.
Published: (2024)
Towards Scalable Foundation Model for Multi-modal and Hyperspectral Geospatial Data
by: Si, Haozhe, et al.
Published: (2025)
by: Si, Haozhe, et al.
Published: (2025)
POEM: Precise Object-level Editing via MLLM control
by: Schouten, Marco, et al.
Published: (2025)
by: Schouten, Marco, et al.
Published: (2025)
Learning Long-Term Temporal Dependencies in Photovoltaic Power Output Prediction Through Multi-Horizon Forecasting
by: Laha, Sumit, et al.
Published: (2026)
by: Laha, Sumit, et al.
Published: (2026)
Pretext Task Adversarial Learning for Unpaired Low-field to Ultra High-field MRI Synthesis
by: Zhang, Zhenxuan, et al.
Published: (2025)
by: Zhang, Zhenxuan, et al.
Published: (2025)
All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment
by: Zhang, Chunhui, et al.
Published: (2023)
by: Zhang, Chunhui, et al.
Published: (2023)
AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval
by: Wang, Yihan, et al.
Published: (2026)
by: Wang, Yihan, et al.
Published: (2026)
Learning Progressive Adaptation for Multi-Modal Tracking
by: Wang, He, et al.
Published: (2026)
by: Wang, He, et al.
Published: (2026)
Similar Items
-
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
by: Gordon, Lucia, et al.
Published: (2026) -
Nacala-Roof-Material: Drone Imagery for Roof Detection, Classification, and Segmentation to Support Mosquito-borne Disease Risk Assessment
by: Guthula, Venkanna Babu, et al.
Published: (2024) -
Familiarity-Based Open-Set Recognition Under Adversarial Attacks
by: Enevoldsen, Philip, et al.
Published: (2023) -
SuperF: Neural Implicit Fields for Multi-Image Super-Resolution
by: Jyhne, Sander Riisøen, et al.
Published: (2025) -
Labeled Data Selection for Category Discovery
by: Zhao, Bingchen, et al.
Published: (2024)