Better to Ask in English: Evaluation of Large Language Models on English, Low-resource and Cross-Lingual Settings

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dey, Krishno, Tarannum, Prerona, Hasan, Md. Arid, Razzak, Imran, Naseem, Usman
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914976083476480
author Dey, Krishno
Tarannum, Prerona
Hasan, Md. Arid
Razzak, Imran
Naseem, Usman
author_facet Dey, Krishno
Tarannum, Prerona
Hasan, Md. Arid
Razzak, Imran
Naseem, Usman
contents Large Language Models (LLMs) are trained on massive amounts of data, enabling their application across diverse domains and tasks. Despite their remarkable performance, most LLMs are developed and evaluated primarily in English. Recently, a few multi-lingual LLMs have emerged, but their performance in low-resource languages, especially the most spoken languages in South Asia, is less explored. To address this gap, in this study, we evaluate LLMs such as GPT-4, Llama 2, and Gemini to analyze their effectiveness in English compared to other low-resource languages from South Asia (e.g., Bangla, Hindi, and Urdu). Specifically, we utilized zero-shot prompting and five different prompt settings to extensively investigate the effectiveness of the LLMs in cross-lingual translated prompts. The findings of the study suggest that GPT-4 outperformed Llama 2 and Gemini in all five prompt settings and across all languages. Moreover, all three LLMs performed better for English language prompts than other low-resource language prompts. This study extensively investigates LLMs in low-resource language contexts to highlight the improvements required in LLMs and language-specific resources to develop more generally purposed NLP applications.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13153
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Better to Ask in English: Evaluation of Large Language Models on English, Low-resource and Cross-Lingual Settings
Dey, Krishno
Tarannum, Prerona
Hasan, Md. Arid
Razzak, Imran
Naseem, Usman
Computation and Language
Large Language Models (LLMs) are trained on massive amounts of data, enabling their application across diverse domains and tasks. Despite their remarkable performance, most LLMs are developed and evaluated primarily in English. Recently, a few multi-lingual LLMs have emerged, but their performance in low-resource languages, especially the most spoken languages in South Asia, is less explored. To address this gap, in this study, we evaluate LLMs such as GPT-4, Llama 2, and Gemini to analyze their effectiveness in English compared to other low-resource languages from South Asia (e.g., Bangla, Hindi, and Urdu). Specifically, we utilized zero-shot prompting and five different prompt settings to extensively investigate the effectiveness of the LLMs in cross-lingual translated prompts. The findings of the study suggest that GPT-4 outperformed Llama 2 and Gemini in all five prompt settings and across all languages. Moreover, all three LLMs performed better for English language prompts than other low-resource language prompts. This study extensively investigates LLMs in low-resource language contexts to highlight the improvements required in LLMs and language-specific resources to develop more generally purposed NLP applications.
title Better to Ask in English: Evaluation of Large Language Models on English, Low-resource and Cross-Lingual Settings
topic Computation and Language
url https://arxiv.org/abs/2410.13153