Questions beyond Pixels: Integrating Commonsense Knowledge in Visual Question Generation for Remote Sensing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Siran, Mi, Li, Castillo-Navarro, Javiera, Tuia, Devis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Knowledge-aware Visual Question Generation for Remote Sensing Images
von: Li, Siran, et al.
Veröffentlicht: (2026)
von: Li, Siran, et al.
Veröffentlicht: (2026)
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
von: Mi, Li, et al.
Veröffentlicht: (2024)
von: Mi, Li, et al.
Veröffentlicht: (2024)
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
von: Mi, Li, et al.
Veröffentlicht: (2024)
von: Mi, Li, et al.
Veröffentlicht: (2024)
ConGeo: Robust Cross-view Geo-localization across Ground View Variations
von: Mi, Li, et al.
Veröffentlicht: (2024)
von: Mi, Li, et al.
Veröffentlicht: (2024)
EcoWikiRS: Learning Ecological Representation of Satellite Images from Weak Supervision with Species Observations and Wikipedia
von: Zermatten, Valerie, et al.
Veröffentlicht: (2025)
von: Zermatten, Valerie, et al.
Veröffentlicht: (2025)
SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
von: Sumbul, Gencer, et al.
Veröffentlicht: (2025)
von: Sumbul, Gencer, et al.
Veröffentlicht: (2025)
Multilingual Vision-Language Pre-training for the Remote Sensing Domain
von: Silva, João Daniel, et al.
Veröffentlicht: (2024)
von: Silva, João Daniel, et al.
Veröffentlicht: (2024)
Large Language Models for Captioning and Retrieving Remote Sensing Images
von: Silva, João Daniel, et al.
Veröffentlicht: (2024)
von: Silva, João Daniel, et al.
Veröffentlicht: (2024)
An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
von: Silva, João Daniel, et al.
Veröffentlicht: (2025)
von: Silva, João Daniel, et al.
Veröffentlicht: (2025)
What to align in multimodal contrastive learning?
von: Dufumier, Benoit, et al.
Veröffentlicht: (2024)
von: Dufumier, Benoit, et al.
Veröffentlicht: (2024)
GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
von: Mi, Li, et al.
Veröffentlicht: (2025)
von: Mi, Li, et al.
Veröffentlicht: (2025)
Visual Question Answering on Multiple Remote Sensing Image Modalities
von: Boussaid, Hichem, et al.
Veröffentlicht: (2025)
von: Boussaid, Hichem, et al.
Veröffentlicht: (2025)
Text-Guided Coarse-to-Fine Fusion Network for Robust Remote Sensing Visual Question Answering
von: Zhao, Zhicheng, et al.
Veröffentlicht: (2024)
von: Zhao, Zhicheng, et al.
Veröffentlicht: (2024)
Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images
von: Tosato, Lucrezia, et al.
Veröffentlicht: (2024)
von: Tosato, Lucrezia, et al.
Veröffentlicht: (2024)
Large Vision-Language Models for Remote Sensing Visual Question Answering
von: Siripong, Surasakdi, et al.
Veröffentlicht: (2024)
von: Siripong, Surasakdi, et al.
Veröffentlicht: (2024)
Cross-Modal Learning of Housing Quality in Amsterdam
von: Levering, Alex, et al.
Veröffentlicht: (2024)
von: Levering, Alex, et al.
Veröffentlicht: (2024)
POLO -- Point-based, multi-class animal detection
von: May, Giacomo, et al.
Veröffentlicht: (2024)
von: May, Giacomo, et al.
Veröffentlicht: (2024)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
von: Wang, Yuduo, et al.
Veröffentlicht: (2023)
von: Wang, Yuduo, et al.
Veröffentlicht: (2023)
SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering
von: Park, Jiho, et al.
Veröffentlicht: (2026)
von: Park, Jiho, et al.
Veröffentlicht: (2026)
Retrieval of Surface Solar Radiation through Implicit Albedo Recovery from Temporal Context
von: Frischholz, Yael, et al.
Veröffentlicht: (2025)
von: Frischholz, Yael, et al.
Veröffentlicht: (2025)
High-resolution Population Maps Derived from Sentinel-1 and Sentinel-2
von: Metzger, Nando, et al.
Veröffentlicht: (2023)
von: Metzger, Nando, et al.
Veröffentlicht: (2023)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
Show Me What and Where has Changed? Question Answering and Grounding for Remote Sensing Change Detection
von: Li, Ke, et al.
Veröffentlicht: (2024)
von: Li, Ke, et al.
Veröffentlicht: (2024)
Augmented Commonsense Knowledge for Remote Object Grounding
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2024)
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2024)
Multi-Scale Grouped Prototypes for Interpretable Semantic Segmentation
von: Porta, Hugo, et al.
Veröffentlicht: (2024)
von: Porta, Hugo, et al.
Veröffentlicht: (2024)
CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities
von: Porta, Hugo, et al.
Veröffentlicht: (2025)
von: Porta, Hugo, et al.
Veröffentlicht: (2025)
Structure Causal Models and LLMs Integration in Medical Visual Question Answering
von: Xu, Zibo, et al.
Veröffentlicht: (2025)
von: Xu, Zibo, et al.
Veröffentlicht: (2025)
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
von: Zi, Xing, et al.
Veröffentlicht: (2025)
von: Zi, Xing, et al.
Veröffentlicht: (2025)
Copy-Move Forgery Detection and Question Answering for Remote Sensing Image
von: Zhang, Ze, et al.
Veröffentlicht: (2024)
von: Zhang, Ze, et al.
Veröffentlicht: (2024)
Object Retrieval for Visual Question Answering with Outside Knowledge
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
Knowledge Detection by Relevant Question and Image Attributes in Visual Question Answering
von: Ahir, Param, et al.
Veröffentlicht: (2023)
von: Ahir, Param, et al.
Veröffentlicht: (2023)
Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
von: Oh, Ju-Young
Veröffentlicht: (2025)
von: Oh, Ju-Young
Veröffentlicht: (2025)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
von: Litrico, Mattia, et al.
Veröffentlicht: (2025)
von: Litrico, Mattia, et al.
Veröffentlicht: (2025)
Acknowledging Focus Ambiguity in Visual Questions
von: Chen, Chongyan, et al.
Veröffentlicht: (2025)
von: Chen, Chongyan, et al.
Veröffentlicht: (2025)
Multi-Sourced Compositional Generalization in Visual Question Answering
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
von: Li, Zhiyang, et al.
Veröffentlicht: (2026)
von: Li, Zhiyang, et al.
Veröffentlicht: (2026)
Multimodal Rationales for Explainable Visual Question Answering
von: Li, Kun, et al.
Veröffentlicht: (2024)
von: Li, Kun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Knowledge-aware Visual Question Generation for Remote Sensing Images
von: Li, Siran, et al.
Veröffentlicht: (2026) -
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
von: Mi, Li, et al.
Veröffentlicht: (2024) -
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
von: Mi, Li, et al.
Veröffentlicht: (2024) -
ConGeo: Robust Cross-view Geo-localization across Ground View Variations
von: Mi, Li, et al.
Veröffentlicht: (2024) -
EcoWikiRS: Learning Ecological Representation of Satellite Images from Weak Supervision with Species Observations and Wikipedia
von: Zermatten, Valerie, et al.
Veröffentlicht: (2025)