VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Islam, Md. Adnanul, Sayeedi, Md. Faiyaz Abdullah, Shuvo, Md. Asaduzzaman, Bappy, Shahanur Rahman, Islam, Md Asiful, Shatabda, Swakkhar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915932588212224
author Islam, Md. Adnanul
Sayeedi, Md. Faiyaz Abdullah
Shuvo, Md. Asaduzzaman
Bappy, Shahanur Rahman
Islam, Md Asiful
Shatabda, Swakkhar
author_facet Islam, Md. Adnanul
Sayeedi, Md. Faiyaz Abdullah
Shuvo, Md. Asaduzzaman
Bappy, Shahanur Rahman
Islam, Md Asiful
Shatabda, Swakkhar
contents Mosquito-borne diseases pose a major global health risk, requiring early detection and proactive control of breeding sites to prevent outbreaks. In this paper, we present VisText-Mosquito, a multimodal dataset that integrates visual and textual data to support automated detection, segmentation, and explanation for mosquito breeding site analysis. The dataset includes 1,828 annotated images for object detection, 142 images for water surface segmentation, and natural language explanation texts linked to each image. The YOLOv9s model achieves the highest precision of 0.92926 and mAP@50 of 0.92891 for object detection, while YOLOv11n-Seg reaches a segmentation precision of 0.91587 and mAP@50 of 0.79795. For textual explanation generation, we tested a range of large vision-language models (LVLMs) in both zero-shot and few-shot settings. Our fine-tuned Mosquito-LLaMA3-8B model achieved the best results, with a final loss of 0.0028, a BLEU score of 54.7, BERTScore of 0.91, and ROUGE-L of 0.85. This dataset and model framework emphasize the theme "Prevention is Better than Cure", showcasing how AI-based detection can proactively address mosquito-borne disease risks. The dataset and implementation code are publicly available at GitHub: https://github.com/adnanul-islam-jisun/VisText-Mosquito
format Preprint
id arxiv_https___arxiv_org_abs_2506_14629
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites
Islam, Md. Adnanul
Sayeedi, Md. Faiyaz Abdullah
Shuvo, Md. Asaduzzaman
Bappy, Shahanur Rahman
Islam, Md Asiful
Shatabda, Swakkhar
Computer Vision and Pattern Recognition
Computation and Language
Mosquito-borne diseases pose a major global health risk, requiring early detection and proactive control of breeding sites to prevent outbreaks. In this paper, we present VisText-Mosquito, a multimodal dataset that integrates visual and textual data to support automated detection, segmentation, and explanation for mosquito breeding site analysis. The dataset includes 1,828 annotated images for object detection, 142 images for water surface segmentation, and natural language explanation texts linked to each image. The YOLOv9s model achieves the highest precision of 0.92926 and mAP@50 of 0.92891 for object detection, while YOLOv11n-Seg reaches a segmentation precision of 0.91587 and mAP@50 of 0.79795. For textual explanation generation, we tested a range of large vision-language models (LVLMs) in both zero-shot and few-shot settings. Our fine-tuned Mosquito-LLaMA3-8B model achieved the best results, with a final loss of 0.0028, a BLEU score of 54.7, BERTScore of 0.91, and ROUGE-L of 0.85. This dataset and model framework emphasize the theme "Prevention is Better than Cure", showcasing how AI-based detection can proactively address mosquito-borne disease risks. The dataset and implementation code are publicly available at GitHub: https://github.com/adnanul-islam-jisun/VisText-Mosquito
title VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2506.14629