RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yuduo, Ghamisi, Pedram |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Change Captioning in Remote Sensing: Evolution to SAT-Cap -- A Single-Stage Transformer Approach
by: Wang, Yuduo, et al.
Published: (2025)
by: Wang, Yuduo, et al.
Published: (2025)
ChangeMinds: Multi-task Framework for Detecting and Describing Changes in Remote Sensing
by: Wang, Yuduo, et al.
Published: (2024)
by: Wang, Yuduo, et al.
Published: (2024)
Mask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change Captioning
by: Sun, Dongwei, et al.
Published: (2024)
by: Sun, Dongwei, et al.
Published: (2024)
FlexiMo: A Flexible Remote Sensing Foundation Model
by: Li, Xuyang, et al.
Published: (2025)
by: Li, Xuyang, et al.
Published: (2025)
Towards Realistic Remote Sensing Dataset Distillation with Discriminative Prototype-guided Diffusion
by: Xu, Yonghao, et al.
Published: (2026)
by: Xu, Yonghao, et al.
Published: (2026)
Questioning the Stability of Visual Question Answering
by: Rosenfeld, Amir, et al.
Published: (2025)
by: Rosenfeld, Amir, et al.
Published: (2025)
Universal Adversarial Defense in Remote Sensing Based on Pre-trained Denoising Diffusion Models
by: Yu, Weikang, et al.
Published: (2023)
by: Yu, Weikang, et al.
Published: (2023)
SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering
by: Li, Bingxin
Published: (2025)
by: Li, Bingxin
Published: (2025)
Diffusion Models Meet Remote Sensing: Principles, Methods, and Perspectives
by: Liu, Yidan, et al.
Published: (2024)
by: Liu, Yidan, et al.
Published: (2024)
Describe Anything Model for Visual Question Answering on Text-rich Images
by: Vu, Yen-Linh, et al.
Published: (2025)
by: Vu, Yen-Linh, et al.
Published: (2025)
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion
by: Chen, Peiyuan, et al.
Published: (2024)
by: Chen, Peiyuan, et al.
Published: (2024)
BERT-VQA: Visual Question Answering on Plots
by: Vu, Tai, et al.
Published: (2025)
by: Vu, Tai, et al.
Published: (2025)
Promptable Foundation Models for SAR Remote Sensing: Adapting the Segment Anything Model for Snow Avalanche Segmentation
by: Gelato, Riccardo, et al.
Published: (2026)
by: Gelato, Riccardo, et al.
Published: (2026)
Spatial Gated Multi-Layer Perceptron for Land Use and Land Cover Mapping
by: Jamali, Ali, et al.
Published: (2023)
by: Jamali, Ali, et al.
Published: (2023)
TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
by: Akl, Ahmed, et al.
Published: (2024)
by: Akl, Ahmed, et al.
Published: (2024)
CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering
by: Huai, Tianyu, et al.
Published: (2025)
by: Huai, Tianyu, et al.
Published: (2025)
Geospatial Foundation Models to Enable Progress on Sustainable Development Goals
by: Ghamisi, Pedram, et al.
Published: (2025)
by: Ghamisi, Pedram, et al.
Published: (2025)
Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
by: Hartsock, Iryna, et al.
Published: (2024)
by: Hartsock, Iryna, et al.
Published: (2024)
MaskCD: A Remote Sensing Change Detection Network Based on Mask Classification
by: Yu, Weikang, et al.
Published: (2024)
by: Yu, Weikang, et al.
Published: (2024)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
by: Singh, Shubhankar, et al.
Published: (2024)
by: Singh, Shubhankar, et al.
Published: (2024)
SeaMo: A Season-Aware Multimodal Foundation Model for Remote Sensing
by: Li, Xuyang, et al.
Published: (2024)
by: Li, Xuyang, et al.
Published: (2024)
FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing
by: Corley, Isaac, et al.
Published: (2025)
by: Corley, Isaac, et al.
Published: (2025)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering
by: Weng, Weixi, et al.
Published: (2024)
by: Weng, Weixi, et al.
Published: (2024)
MineNetCD: A Benchmark for Global Mining Change Detection on Remote Sensing Imagery
by: Yu, Weikang, et al.
Published: (2024)
by: Yu, Weikang, et al.
Published: (2024)
Visual Question Answering on Multiple Remote Sensing Image Modalities
by: Boussaid, Hichem, et al.
Published: (2025)
by: Boussaid, Hichem, et al.
Published: (2025)
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
by: Mamaghan, Amir Mohammad Karimi, et al.
Published: (2024)
by: Mamaghan, Amir Mohammad Karimi, et al.
Published: (2024)
Score-Based Multimodal Autoencoder
by: Wesego, Daniel, et al.
Published: (2023)
by: Wesego, Daniel, et al.
Published: (2023)
Multimodal ELBO with Diffusion Decoders
by: Wesego, Daniel, et al.
Published: (2024)
by: Wesego, Daniel, et al.
Published: (2024)
Large Vision-Language Models for Remote Sensing Visual Question Answering
by: Siripong, Surasakdi, et al.
Published: (2024)
by: Siripong, Surasakdi, et al.
Published: (2024)
Towards Knowledge Guided Pretraining Approaches for Multimodal Foundation Models: Applications in Remote Sensing
by: Ravirathinam, Praveen, et al.
Published: (2024)
by: Ravirathinam, Praveen, et al.
Published: (2024)
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
by: Indrehus, Kjetil, et al.
Published: (2026)
by: Indrehus, Kjetil, et al.
Published: (2026)
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
by: Chintapatla, Ishant, et al.
Published: (2025)
by: Chintapatla, Ishant, et al.
Published: (2025)
Privacy-Aware Document Visual Question Answering
by: Tito, Rubèn, et al.
Published: (2023)
by: Tito, Rubèn, et al.
Published: (2023)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering
by: Yu, Zhou, et al.
Published: (2023)
by: Yu, Zhou, et al.
Published: (2023)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
by: Yanuka, Moran, et al.
Published: (2024)
by: Yanuka, Moran, et al.
Published: (2024)
TinyVQA: Compact Multimodal Deep Neural Network for Visual Question Answering on Resource-Constrained Devices
by: Rashid, Hasib-Al, et al.
Published: (2024)
by: Rashid, Hasib-Al, et al.
Published: (2024)
Vision Foundation Models in Remote Sensing: A Survey
by: Lu, Siqi, et al.
Published: (2024)
by: Lu, Siqi, et al.
Published: (2024)
Exploring Diverse Methods in Visual Question Answering
by: Li, Panfeng, et al.
Published: (2024)
by: Li, Panfeng, et al.
Published: (2024)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
by: Patel, Alkesh, et al.
Published: (2025)
by: Patel, Alkesh, et al.
Published: (2025)
Similar Items
-
Change Captioning in Remote Sensing: Evolution to SAT-Cap -- A Single-Stage Transformer Approach
by: Wang, Yuduo, et al.
Published: (2025) -
ChangeMinds: Multi-task Framework for Detecting and Describing Changes in Remote Sensing
by: Wang, Yuduo, et al.
Published: (2024) -
Mask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change Captioning
by: Sun, Dongwei, et al.
Published: (2024) -
FlexiMo: A Flexible Remote Sensing Foundation Model
by: Li, Xuyang, et al.
Published: (2025) -
Towards Realistic Remote Sensing Dataset Distillation with Discriminative Prototype-guided Diffusion
by: Xu, Yonghao, et al.
Published: (2026)