RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation
Fuente:
arXiv
Saved in:
| Main Authors: | Houdré, Nicolas, Marcos, Diego, de Turckheim, Hugo Riffaud, Ienco, Dino, Wendling, Laurent, Kurtz, Camille, Lobry, Sylvain |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Atomizer: Generalizing to new modalities by breaking satellite images down to a set of scalars
by: de Turckheim, Hugo Riffaud, et al.
Published: (2025)
by: de Turckheim, Hugo Riffaud, et al.
Published: (2025)
Visual Question Answering on Multiple Remote Sensing Image Modalities
by: Boussaid, Hichem, et al.
Published: (2025)
by: Boussaid, Hichem, et al.
Published: (2025)
Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images
by: Tosato, Lucrezia, et al.
Published: (2024)
by: Tosato, Lucrezia, et al.
Published: (2024)
Can SAR improve RSVQA performance?
by: Tosato, Lucrezia, et al.
Published: (2024)
by: Tosato, Lucrezia, et al.
Published: (2024)
SAR Strikes Back: A New Hope for RSVQA
by: Tosato, Lucrezia, et al.
Published: (2025)
by: Tosato, Lucrezia, et al.
Published: (2025)
SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting
by: Jain, Pallavi, et al.
Published: (2024)
by: Jain, Pallavi, et al.
Published: (2024)
TimeSenCLIP: A Time Series Vision-Language Model for Remote Sensing
by: Jain, Pallavi, et al.
Published: (2025)
by: Jain, Pallavi, et al.
Published: (2025)
Two-stage Vision Transformers and Hard Masking offer Robust Object Representations
by: Aniraj, Ananthu, et al.
Published: (2025)
by: Aniraj, Ananthu, et al.
Published: (2025)
Metonymy in vision models undermines attention-based interpretability
by: Aniraj, Ananthu, et al.
Published: (2026)
by: Aniraj, Ananthu, et al.
Published: (2026)
IC-EO: Interpretable Code-based assistant for Earth Observation
by: Lahouel, Lamia, et al.
Published: (2026)
by: Lahouel, Lamia, et al.
Published: (2026)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025)
by: Mena, Francisco, et al.
Published: (2025)
PDiscoFormer: Relaxing Part Discovery Constraints with Vision Transformers
by: Aniraj, Ananthu, et al.
Published: (2024)
by: Aniraj, Ananthu, et al.
Published: (2024)
MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
by: Labatie, Antoine, et al.
Published: (2025)
by: Labatie, Antoine, et al.
Published: (2025)
DisCoM-KD: Cross-Modal Knowledge Distillation via Disentanglement Representation and Adversarial Learning
by: Ienco, Dino, et al.
Published: (2024)
by: Ienco, Dino, et al.
Published: (2024)
Towards a multimodal framework for remote sensing image change retrieval and captioning
by: Ferrod, Roger, et al.
Published: (2024)
by: Ferrod, Roger, et al.
Published: (2024)
Revisiting Cross-Modal Knowledge Distillation: A Disentanglement Approach for RGBD Semantic Segmentation
by: Ferrod, Roger, et al.
Published: (2025)
by: Ferrod, Roger, et al.
Published: (2025)
AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
by: Astruc, Guillaume, et al.
Published: (2024)
by: Astruc, Guillaume, et al.
Published: (2024)
A Semantically-Aware Relevance Measure for Content-Based Medical Image Retrieval Evaluation
by: Wei, Xiaoyang, et al.
Published: (2025)
by: Wei, Xiaoyang, et al.
Published: (2025)
QwenCLIP: Boosting Medical Vision-Language Pretraining via LLM Embeddings and Prompt tuning
by: Wei, Xiaoyang, et al.
Published: (2025)
by: Wei, Xiaoyang, et al.
Published: (2025)
TerraMesh: A Planetary Mosaic of Multimodal Earth Observation Data
by: Blumenstiel, Benedikt, et al.
Published: (2025)
by: Blumenstiel, Benedikt, et al.
Published: (2025)
UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs
by: Ni, Shuo, et al.
Published: (2026)
by: Ni, Shuo, et al.
Published: (2026)
A Proxy Consistency Loss for Grounded Fusion of Earth Observation and Location Encoders
by: Wang, Zhongying, et al.
Published: (2026)
by: Wang, Zhongying, et al.
Published: (2026)
Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation
by: Xiong, Zhitong, et al.
Published: (2024)
by: Xiong, Zhitong, et al.
Published: (2024)
Checkmate: interpretable and explainable RSVQA is the endgame
by: Tosato, Lucrezia, et al.
Published: (2025)
by: Tosato, Lucrezia, et al.
Published: (2025)
TerraMind: Large-Scale Generative Multimodality for Earth Observation
by: Jakubik, Johannes, et al.
Published: (2025)
by: Jakubik, Johannes, et al.
Published: (2025)
OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
by: Herzog, Henry, et al.
Published: (2025)
by: Herzog, Henry, et al.
Published: (2025)
DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation
by: Xiong, Zhitong, et al.
Published: (2025)
by: Xiong, Zhitong, et al.
Published: (2025)
GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data
by: Ferrod, Roger, et al.
Published: (2026)
by: Ferrod, Roger, et al.
Published: (2026)
Jointly RS Image Deblurring and Super-Resolution with Adjustable-Kernel and Multi-Domain Attention
by: Zhang, Yan, et al.
Published: (2024)
by: Zhang, Yan, et al.
Published: (2024)
OmniSat: Self-Supervised Modality Fusion for Earth Observation
by: Astruc, Guillaume, et al.
Published: (2024)
by: Astruc, Guillaume, et al.
Published: (2024)
SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining
by: Braham, Nassim Ait Ali, et al.
Published: (2026)
by: Braham, Nassim Ait Ali, et al.
Published: (2026)
EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
Deep Multimodal Fusion for Semantic Segmentation of Remote Sensing Earth Observation Data
by: Dimitrovski, Ivica, et al.
Published: (2024)
by: Dimitrovski, Ivica, et al.
Published: (2024)
RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation
by: Min, Rui, et al.
Published: (2026)
by: Min, Rui, et al.
Published: (2026)
EarthNets: Empowering AI in Earth Observation
by: Xiong, Zhitong, et al.
Published: (2022)
by: Xiong, Zhitong, et al.
Published: (2022)
Geographical Context Matters: Bridging Fine and Coarse Spatial Information to Enhance Continental Land Cover Mapping
by: Ghassemi, Babak, et al.
Published: (2025)
by: Ghassemi, Babak, et al.
Published: (2025)
TerraFlow: Multimodal, Multitemporal Representation Learning for Earth Observation
by: Puriy, Nazar, et al.
Published: (2026)
by: Puriy, Nazar, et al.
Published: (2026)
FlowEO: Generative Unsupervised Domain Adaptation for Earth Observation
by: Bellier, Georges Le, et al.
Published: (2025)
by: Bellier, Georges Le, et al.
Published: (2025)
The Change You Want To Detect: Semantic Change Detection In Earth Observation With Hybrid Data Generation
by: Benidir, Yanis, et al.
Published: (2025)
by: Benidir, Yanis, et al.
Published: (2025)
Unified Multimodal Models as Auto-Encoders
by: Yan, Zhiyuan, et al.
Published: (2025)
by: Yan, Zhiyuan, et al.
Published: (2025)
Similar Items
-
Atomizer: Generalizing to new modalities by breaking satellite images down to a set of scalars
by: de Turckheim, Hugo Riffaud, et al.
Published: (2025) -
Visual Question Answering on Multiple Remote Sensing Image Modalities
by: Boussaid, Hichem, et al.
Published: (2025) -
Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images
by: Tosato, Lucrezia, et al.
Published: (2024) -
Can SAR improve RSVQA performance?
by: Tosato, Lucrezia, et al.
Published: (2024) -
SAR Strikes Back: A New Hope for RSVQA
by: Tosato, Lucrezia, et al.
Published: (2025)