Multilingual Vision-Language Pre-training for the Remote Sensing Domain
Fuente:
arXiv
Saved in:
| Main Authors: | Silva, João Daniel, Magalhaes, Joao, Tuia, Devis, Martins, Bruno |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
by: Silva, João Daniel, et al.
Published: (2025)
by: Silva, João Daniel, et al.
Published: (2025)
Large Language Models for Captioning and Retrieving Remote Sensing Images
by: Silva, João Daniel, et al.
Published: (2024)
by: Silva, João Daniel, et al.
Published: (2024)
Multilingual Training-Free Remote Sensing Image Captioning
by: Rebelo, Carlos, et al.
Published: (2025)
by: Rebelo, Carlos, et al.
Published: (2025)
SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
by: Sumbul, Gencer, et al.
Published: (2025)
by: Sumbul, Gencer, et al.
Published: (2025)
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
Knowledge-aware Visual Question Generation for Remote Sensing Images
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
Questions beyond Pixels: Integrating Commonsense Knowledge in Visual Question Generation for Remote Sensing
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
by: Litrico, Mattia, et al.
Published: (2025)
by: Litrico, Mattia, et al.
Published: (2025)
Cross-Modal Learning of Housing Quality in Amsterdam
by: Levering, Alex, et al.
Published: (2024)
by: Levering, Alex, et al.
Published: (2024)
POLO -- Point-based, multi-class animal detection
by: May, Giacomo, et al.
Published: (2024)
by: May, Giacomo, et al.
Published: (2024)
Retrieval of Surface Solar Radiation through Implicit Albedo Recovery from Temporal Context
by: Frischholz, Yael, et al.
Published: (2025)
by: Frischholz, Yael, et al.
Published: (2025)
High-resolution Population Maps Derived from Sentinel-1 and Sentinel-2
by: Metzger, Nando, et al.
Published: (2023)
by: Metzger, Nando, et al.
Published: (2023)
Unsupervised Domain Adaption Harnessing Vision-Language Pre-training
by: Zhou, Wenlve, et al.
Published: (2024)
by: Zhou, Wenlve, et al.
Published: (2024)
Show and Guide: Instructional-Plan Grounded Vision and Language Model
by: Glória-Silva, Diogo, et al.
Published: (2024)
by: Glória-Silva, Diogo, et al.
Published: (2024)
Generic Knowledge Boosted Pre-training For Remote Sensing Images
by: Huang, Ziyue, et al.
Published: (2024)
by: Huang, Ziyue, et al.
Published: (2024)
Multi-Scale Grouped Prototypes for Interpretable Semantic Segmentation
by: Porta, Hugo, et al.
Published: (2024)
by: Porta, Hugo, et al.
Published: (2024)
CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities
by: Porta, Hugo, et al.
Published: (2025)
by: Porta, Hugo, et al.
Published: (2025)
GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
by: Mi, Li, et al.
Published: (2025)
by: Mi, Li, et al.
Published: (2025)
Lightweight, Pre-trained Transformers for Remote Sensing Timeseries
by: Tseng, Gabriel, et al.
Published: (2023)
by: Tseng, Gabriel, et al.
Published: (2023)
EcoWikiRS: Learning Ecological Representation of Satellite Images from Weak Supervision with Species Observations and Wikipedia
by: Zermatten, Valerie, et al.
Published: (2025)
by: Zermatten, Valerie, et al.
Published: (2025)
MutDet: Mutually Optimizing Pre-training for Remote Sensing Object Detection
by: Huang, Ziyue, et al.
Published: (2024)
by: Huang, Ziyue, et al.
Published: (2024)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
by: Condez, Ana Carolina, et al.
Published: (2025)
by: Condez, Ana Carolina, et al.
Published: (2025)
RemoteCLIP: A Vision Language Foundation Model for Remote Sensing
by: Liu, Fan, et al.
Published: (2023)
by: Liu, Fan, et al.
Published: (2023)
Checkmate: interpretable and explainable RSVQA is the endgame
by: Tosato, Lucrezia, et al.
Published: (2025)
by: Tosato, Lucrezia, et al.
Published: (2025)
From Classification to Segmentation with Explainable AI: A Study on Crack Detection and Growth Monitoring
by: Forest, Florent, et al.
Published: (2023)
by: Forest, Florent, et al.
Published: (2023)
Efficient Vision-Language Pre-training by Cluster Masking
by: Wei, Zihao, et al.
Published: (2024)
by: Wei, Zihao, et al.
Published: (2024)
Enhancing Vision-Language Pre-training with Rich Supervisions
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
Towards Privacy-preserved Pre-training of Remote Sensing Foundation Models with Federated Mutual-guidance Learning
by: Tan, Jieyi, et al.
Published: (2025)
by: Tan, Jieyi, et al.
Published: (2025)
Continual Vision-Language Learning for Remote Sensing: Benchmarking and Analysis
by: Weng, Xingxing, et al.
Published: (2026)
by: Weng, Xingxing, et al.
Published: (2026)
Scaling Pre-training to One Hundred Billion Data for Vision Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
VLP: A Survey on Vision-Language Pre-training
by: Chen, Feilong, et al.
Published: (2022)
by: Chen, Feilong, et al.
Published: (2022)
CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models
by: An, Xiao, et al.
Published: (2024)
by: An, Xiao, et al.
Published: (2024)
Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
by: Weng, Xingxing, et al.
Published: (2025)
by: Weng, Xingxing, et al.
Published: (2025)
Redundancy-Aware Pretraining of Vision-Language Foundation Models in Remote Sensing
by: Adler, Mathis Jürgen, et al.
Published: (2025)
by: Adler, Mathis Jürgen, et al.
Published: (2025)
Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing
by: Anderson, Madeline, et al.
Published: (2025)
by: Anderson, Madeline, et al.
Published: (2025)
Few-Shot Adaptation Benchmark for Remote Sensing Vision-Language Models
by: Khoury, Karim El, et al.
Published: (2025)
by: Khoury, Karim El, et al.
Published: (2025)
CrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic Segmentation
by: Gong, Ziyang, et al.
Published: (2024)
by: Gong, Ziyang, et al.
Published: (2024)
ConGeo: Robust Cross-view Geo-localization across Ground View Variations
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
Sample-agnostic Adversarial Perturbation for Vision-Language Pre-training Models
by: Zheng, Haonan, et al.
Published: (2024)
by: Zheng, Haonan, et al.
Published: (2024)
Similar Items
-
An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
by: Silva, João Daniel, et al.
Published: (2025) -
Large Language Models for Captioning and Retrieving Remote Sensing Images
by: Silva, João Daniel, et al.
Published: (2024) -
Multilingual Training-Free Remote Sensing Image Captioning
by: Rebelo, Carlos, et al.
Published: (2025) -
SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
by: Sumbul, Gencer, et al.
Published: (2025) -
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
by: Mi, Li, et al.
Published: (2024)