ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Barua, Deeparghya Dutta, Sourove, Md Sakib Ul Rahman, Fahim, Md, Haider, Fabiha, Shifat, Fariha Tanjim, Adib, Md Tasmim Rahman, Uddin, Anam Borhan, Ishmam, Md Farhan, Alam, Md Farhad
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909630782767104
author Barua, Deeparghya Dutta
Sourove, Md Sakib Ul Rahman
Fahim, Md
Haider, Fabiha
Shifat, Fariha Tanjim
Adib, Md Tasmim Rahman
Uddin, Anam Borhan
Ishmam, Md Farhan
Alam, Md Farhad
author_facet Barua, Deeparghya Dutta
Sourove, Md Sakib Ul Rahman
Fahim, Md
Haider, Fabiha
Shifat, Fariha Tanjim
Adib, Md Tasmim Rahman
Uddin, Anam Borhan
Ishmam, Md Farhan
Alam, Md Farhad
contents Visual Question Answer (VQA) poses the problem of answering a natural language question about a visual context. Bangla, despite being a widely spoken language, is considered low-resource in the realm of VQA due to the lack of proper benchmarks, challenging models known to be performant in other languages. Furthermore, existing Bangla VQA datasets offer little regional relevance and are largely adapted from their foreign counterparts. To address these challenges, we introduce a large-scale Bangla VQA dataset, ChitroJera, totaling over 15k samples from diverse and locally relevant data sources. We assess the performance of text encoders, image encoders, multimodal models, and our novel dual-encoder models. The experiments reveal that the pre-trained dual-encoders outperform other models of their scale. We also evaluate the performance of current large vision language models (LVLMs) using prompt-based techniques, achieving the overall best performance. Given the underdeveloped state of existing datasets, we envision ChitroJera expanding the scope of Vision-Language tasks in Bangla.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14991
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla
Barua, Deeparghya Dutta
Sourove, Md Sakib Ul Rahman
Fahim, Md
Haider, Fabiha
Shifat, Fariha Tanjim
Adib, Md Tasmim Rahman
Uddin, Anam Borhan
Ishmam, Md Farhan
Alam, Md Farhad
Computer Vision and Pattern Recognition
Computation and Language
Visual Question Answer (VQA) poses the problem of answering a natural language question about a visual context. Bangla, despite being a widely spoken language, is considered low-resource in the realm of VQA due to the lack of proper benchmarks, challenging models known to be performant in other languages. Furthermore, existing Bangla VQA datasets offer little regional relevance and are largely adapted from their foreign counterparts. To address these challenges, we introduce a large-scale Bangla VQA dataset, ChitroJera, totaling over 15k samples from diverse and locally relevant data sources. We assess the performance of text encoders, image encoders, multimodal models, and our novel dual-encoder models. The experiments reveal that the pre-trained dual-encoders outperform other models of their scale. We also evaluate the performance of current large vision language models (LVLMs) using prompt-based techniques, achieving the overall best performance. Given the underdeveloped state of existing datasets, we envision ChitroJera expanding the scope of Vision-Language tasks in Bangla.
title ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2410.14991