BEnQA: A Question Answering and Reasoning Benchmark for Bengali and English

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shafayat, Sheikh, Hasan, H M Quamran, Mahim, Minhajur Rahman Chowdhury, Putri, Rifki Afina, Thorne, James, Oh, Alice
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916161444118528
author Shafayat, Sheikh
Hasan, H M Quamran
Mahim, Minhajur Rahman Chowdhury
Putri, Rifki Afina
Thorne, James
Oh, Alice
author_facet Shafayat, Sheikh
Hasan, H M Quamran
Mahim, Minhajur Rahman Chowdhury
Putri, Rifki Afina
Thorne, James
Oh, Alice
contents In this study, we introduce BEnQA, a dataset comprising parallel Bengali and English exam questions for middle and high school levels in Bangladesh. Our dataset consists of approximately 5K questions covering several subjects in science with different types of questions, including factual, application, and reasoning-based questions. We benchmark several Large Language Models (LLMs) with our parallel dataset and observe a notable performance disparity between the models in Bengali and English. We also investigate some prompting methods, and find that Chain-of-Thought prompting is beneficial mostly on reasoning questions, but not so much on factual ones. We also find that appending English translation helps to answer questions in Bengali. Our findings point to promising future research directions for improving the performance of LLMs in Bengali and more generally in low-resource languages.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10900
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BEnQA: A Question Answering and Reasoning Benchmark for Bengali and English
Shafayat, Sheikh
Hasan, H M Quamran
Mahim, Minhajur Rahman Chowdhury
Putri, Rifki Afina
Thorne, James
Oh, Alice
Computation and Language
In this study, we introduce BEnQA, a dataset comprising parallel Bengali and English exam questions for middle and high school levels in Bangladesh. Our dataset consists of approximately 5K questions covering several subjects in science with different types of questions, including factual, application, and reasoning-based questions. We benchmark several Large Language Models (LLMs) with our parallel dataset and observe a notable performance disparity between the models in Bengali and English. We also investigate some prompting methods, and find that Chain-of-Thought prompting is beneficial mostly on reasoning questions, but not so much on factual ones. We also find that appending English translation helps to answer questions in Bengali. Our findings point to promising future research directions for improving the performance of LLMs in Bengali and more generally in low-resource languages.
title BEnQA: A Question Answering and Reasoning Benchmark for Bengali and English
topic Computation and Language
url https://arxiv.org/abs/2403.10900