Visual Question Answering on Multiple Remote Sensing Image Modalities
Fuente:
arXiv
Saved in:
| Main Authors: | Boussaid, Hichem, Tosato, Lucrezia, Weissgerber, Flora, Kurtz, Camille, Wendling, Laurent, Lobry, Sylvain |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images
by: Tosato, Lucrezia, et al.
Published: (2024)
by: Tosato, Lucrezia, et al.
Published: (2024)
SAR Strikes Back: A New Hope for RSVQA
by: Tosato, Lucrezia, et al.
Published: (2025)
by: Tosato, Lucrezia, et al.
Published: (2025)
Can SAR improve RSVQA performance?
by: Tosato, Lucrezia, et al.
Published: (2024)
by: Tosato, Lucrezia, et al.
Published: (2024)
Checkmate: interpretable and explainable RSVQA is the endgame
by: Tosato, Lucrezia, et al.
Published: (2025)
by: Tosato, Lucrezia, et al.
Published: (2025)
Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning
by: Tosato, Lucrezia, et al.
Published: (2026)
by: Tosato, Lucrezia, et al.
Published: (2026)
RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation
by: Houdré, Nicolas, et al.
Published: (2025)
by: Houdré, Nicolas, et al.
Published: (2025)
Large Vision-Language Models for Remote Sensing Visual Question Answering
by: Siripong, Surasakdi, et al.
Published: (2024)
by: Siripong, Surasakdi, et al.
Published: (2024)
Copy-Move Forgery Detection and Question Answering for Remote Sensing Image
by: Zhang, Ze, et al.
Published: (2024)
by: Zhang, Ze, et al.
Published: (2024)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
by: Wang, Yuduo, et al.
Published: (2023)
by: Wang, Yuduo, et al.
Published: (2023)
Knowledge-aware Visual Question Generation for Remote Sensing Images
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
Text-Guided Coarse-to-Fine Fusion Network for Robust Remote Sensing Visual Question Answering
by: Zhao, Zhicheng, et al.
Published: (2024)
by: Zhao, Zhicheng, et al.
Published: (2024)
Exploiting temporal information to detect conversational groups in videos and predict the next speaker
by: Tosato, Lucrezia, et al.
Published: (2024)
by: Tosato, Lucrezia, et al.
Published: (2024)
RS-MoE: A Vision-Language Model with Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering
by: Lin, Hui, et al.
Published: (2024)
by: Lin, Hui, et al.
Published: (2024)
Questions beyond Pixels: Integrating Commonsense Knowledge in Visual Question Generation for Remote Sensing
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
A Semantically-Aware Relevance Measure for Content-Based Medical Image Retrieval Evaluation
by: Wei, Xiaoyang, et al.
Published: (2025)
by: Wei, Xiaoyang, et al.
Published: (2025)
Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation
by: Nikandrou, Malvina, et al.
Published: (2024)
by: Nikandrou, Malvina, et al.
Published: (2024)
RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
by: Zi, Xing, et al.
Published: (2025)
by: Zi, Xing, et al.
Published: (2025)
Show Me What and Where has Changed? Question Answering and Grounding for Remote Sensing Change Detection
by: Li, Ke, et al.
Published: (2024)
by: Li, Ke, et al.
Published: (2024)
QIRL: Boosting Visual Question Answering via Optimized Question-Image Relation Learning
by: Xu, Quanxing, et al.
Published: (2025)
by: Xu, Quanxing, et al.
Published: (2025)
Selectively Answering Visual Questions
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
Visually Interpretable Subtask Reasoning for Visual Question Answering
by: Cheng, Yu, et al.
Published: (2025)
by: Cheng, Yu, et al.
Published: (2025)
Targeted Visual Prompting for Medical Visual Question Answering
by: Tascon-Morales, Sergio, et al.
Published: (2024)
by: Tascon-Morales, Sergio, et al.
Published: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
Knowledge Detection by Relevant Question and Image Attributes in Visual Question Answering
by: Ahir, Param, et al.
Published: (2023)
by: Ahir, Param, et al.
Published: (2023)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
by: Kim, Hongyeob, et al.
Published: (2025)
by: Kim, Hongyeob, et al.
Published: (2025)
Questioning the Stability of Visual Question Answering
by: Rosenfeld, Amir, et al.
Published: (2025)
by: Rosenfeld, Amir, et al.
Published: (2025)
MedCFVQA: A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering
by: Ye, Shuchang, et al.
Published: (2025)
by: Ye, Shuchang, et al.
Published: (2025)
A Multi-Modal Federated Learning Framework for Remote Sensing Image Classification
by: Büyüktaş, Barış, et al.
Published: (2025)
by: Büyüktaş, Barış, et al.
Published: (2025)
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
by: Shen, Zhixuan, et al.
Published: (2024)
by: Shen, Zhixuan, et al.
Published: (2024)
Evaluating Variance in Visual Question Answering Benchmarks
by: SR, Nikitha
Published: (2025)
by: SR, Nikitha
Published: (2025)
Multimodal Rationales for Explainable Visual Question Answering
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Atomizer: Generalizing to new modalities by breaking satellite images down to a set of scalars
by: de Turckheim, Hugo Riffaud, et al.
Published: (2025)
by: de Turckheim, Hugo Riffaud, et al.
Published: (2025)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
by: Park, Kyu Ri, et al.
Published: (2024)
by: Park, Kyu Ri, et al.
Published: (2024)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
by: Özdemir, Övgü, et al.
Published: (2024)
by: Özdemir, Övgü, et al.
Published: (2024)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024)
by: Romero, David, et al.
Published: (2024)
QwenCLIP: Boosting Medical Vision-Language Pretraining via LLM Embeddings and Prompt tuning
by: Wei, Xiaoyang, et al.
Published: (2025)
by: Wei, Xiaoyang, et al.
Published: (2025)
Image augmentation with invertible networks in interactive satellite image change detection
by: Sahbi, Hichem
Published: (2025)
by: Sahbi, Hichem
Published: (2025)
Self-Supervised Cross-Modal Text-Image Time Series Retrieval in Remote Sensing
by: Hoxha, Genc, et al.
Published: (2025)
by: Hoxha, Genc, et al.
Published: (2025)
DriveLM: Driving with Graph Visual Question Answering
by: Sima, Chonghao, et al.
Published: (2023)
by: Sima, Chonghao, et al.
Published: (2023)
Object Retrieval for Visual Question Answering with Outside Knowledge
by: Kan, Shichao, et al.
Published: (2024)
by: Kan, Shichao, et al.
Published: (2024)
Similar Items
-
Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images
by: Tosato, Lucrezia, et al.
Published: (2024) -
SAR Strikes Back: A New Hope for RSVQA
by: Tosato, Lucrezia, et al.
Published: (2025) -
Can SAR improve RSVQA performance?
by: Tosato, Lucrezia, et al.
Published: (2024) -
Checkmate: interpretable and explainable RSVQA is the endgame
by: Tosato, Lucrezia, et al.
Published: (2025) -
Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning
by: Tosato, Lucrezia, et al.
Published: (2026)