Multilingual Performance Biases of Large Language Models in Education

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Vansh, Chowdhury, Sankalan Pal, Zouhar, Vilém, Rooein, Donya, Sachan, Mrinmaya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908479012208640
author Gupta, Vansh
Chowdhury, Sankalan Pal
Zouhar, Vilém
Rooein, Donya
Sachan, Mrinmaya
author_facet Gupta, Vansh
Chowdhury, Sankalan Pal
Zouhar, Vilém
Rooein, Donya
Sachan, Mrinmaya
contents Large language models (LLMs) are increasingly being adopted in educational settings. These applications expand beyond English, though current LLMs remain primarily English-centric. In this work, we ascertain if their use in education settings in non-English languages is warranted. We evaluated the performance of popular LLMs on four educational tasks: identifying student misconceptions, providing targeted feedback, interactive tutoring, and grading translations in eight languages (Mandarin, Hindi, Arabic, German, Farsi, Telugu, Ukrainian, Czech) in addition to English. We find that the performance on these tasks somewhat corresponds to the amount of language represented in training data, with lower-resource languages having poorer task performance. Although the models perform reasonably well in most languages, the frequent performance drop from English is significant. Thus, we recommend that practitioners first verify that the LLM works well in the target language for their educational task before deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17720
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multilingual Performance Biases of Large Language Models in Education
Gupta, Vansh
Chowdhury, Sankalan Pal
Zouhar, Vilém
Rooein, Donya
Sachan, Mrinmaya
Computation and Language
Artificial Intelligence
Large language models (LLMs) are increasingly being adopted in educational settings. These applications expand beyond English, though current LLMs remain primarily English-centric. In this work, we ascertain if their use in education settings in non-English languages is warranted. We evaluated the performance of popular LLMs on four educational tasks: identifying student misconceptions, providing targeted feedback, interactive tutoring, and grading translations in eight languages (Mandarin, Hindi, Arabic, German, Farsi, Telugu, Ukrainian, Czech) in addition to English. We find that the performance on these tasks somewhat corresponds to the amount of language represented in training data, with lower-resource languages having poorer task performance. Although the models perform reasonably well in most languages, the frequent performance drop from English is significant. Thus, we recommend that practitioners first verify that the LLM works well in the target language for their educational task before deployment.
title Multilingual Performance Biases of Large Language Models in Education
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.17720