DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Al-Mohannadi, Aisha, Firoz, Ayisha, Yang, Yin, Imran, Muhammad, Ofli, Ferda
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909053549019136
author Al-Mohannadi, Aisha
Firoz, Ayisha
Yang, Yin
Imran, Muhammad
Ofli, Ferda
author_facet Al-Mohannadi, Aisha
Firoz, Ayisha
Yang, Yin
Imran, Muhammad
Ofli, Ferda
contents Social media imagery provides a low-latency source of situational information during natural and human-induced disasters, enabling rapid damage assessment and response. While Visual Question Answering (VQA) has shown strong performance in general-purpose domains, its suitability for the complex and safety-critical reasoning required in disaster response remains unclear. We introduce DisasterVQA, a benchmark dataset designed for perception and reasoning in crisis contexts. DisasterVQA consists of 1,395 real-world images and 4,405 expert-curated question-answer pairs spanning diverse events such as floods, wildfires, and earthquakes. Grounded in humanitarian frameworks including FEMA ESF and OCHA MIRA, the dataset includes binary, multiple-choice, and open-ended questions covering situational awareness and operational decision-making tasks. We benchmark seven state-of-the-art vision-language models and find performance variability across question types, disaster categories, regions, and humanitarian tasks. Although models achieve high accuracy on binary questions, they struggle with fine-grained quantitative reasoning, object counting, and context-sensitive interpretation, particularly for underrepresented disaster scenarios. DisasterVQA provides a challenging and practical benchmark to guide the development of more robust and operationally meaningful vision-language models for disaster response. The dataset is publicly available at https://doi.org/10.5281/zenodo.18267769.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13839
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
Al-Mohannadi, Aisha
Firoz, Ayisha
Yang, Yin
Imran, Muhammad
Ofli, Ferda
Computer Vision and Pattern Recognition
Social media imagery provides a low-latency source of situational information during natural and human-induced disasters, enabling rapid damage assessment and response. While Visual Question Answering (VQA) has shown strong performance in general-purpose domains, its suitability for the complex and safety-critical reasoning required in disaster response remains unclear. We introduce DisasterVQA, a benchmark dataset designed for perception and reasoning in crisis contexts. DisasterVQA consists of 1,395 real-world images and 4,405 expert-curated question-answer pairs spanning diverse events such as floods, wildfires, and earthquakes. Grounded in humanitarian frameworks including FEMA ESF and OCHA MIRA, the dataset includes binary, multiple-choice, and open-ended questions covering situational awareness and operational decision-making tasks. We benchmark seven state-of-the-art vision-language models and find performance variability across question types, disaster categories, regions, and humanitarian tasks. Although models achieve high accuracy on binary questions, they struggle with fine-grained quantitative reasoning, object counting, and context-sensitive interpretation, particularly for underrepresented disaster scenarios. DisasterVQA provides a challenging and practical benchmark to guide the development of more robust and operationally meaningful vision-language models for disaster response. The dataset is publicly available at https://doi.org/10.5281/zenodo.18267769.
title DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.13839