On the robustness of ChatGPT in teaching Korean Mathematics

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nguyen, Phuong-Nam, Nguyen-The, Quang, Vu-Minh, An, Nguyen, Diep-Anh, Pham, Xuan-Lam
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929718243098624
author Nguyen, Phuong-Nam
Nguyen-The, Quang
Vu-Minh, An
Nguyen, Diep-Anh
Pham, Xuan-Lam
author_facet Nguyen, Phuong-Nam
Nguyen-The, Quang
Vu-Minh, An
Nguyen, Diep-Anh
Pham, Xuan-Lam
contents ChatGPT, an Artificial Intelligence model, has the potential to revolutionize education. However, its effectiveness in solving non-English questions remains uncertain. This study evaluates ChatGPT's robustness using 586 Korean mathematics questions. ChatGPT achieves 66.72% accuracy, correctly answering 391 out of 586 questions. We also assess its ability to rate mathematics questions based on eleven criteria and perform a topic analysis. Our findings show that ChatGPT's ratings align with educational theory and test-taker perspectives. While ChatGPT performs well in question classification, it struggles with non-English contexts, highlighting areas for improvement. Future research should address linguistic biases and enhance accuracy across diverse languages. Domain-specific optimizations and multilingual training could improve ChatGPT's role in personalized education.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11915
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the robustness of ChatGPT in teaching Korean Mathematics
Nguyen, Phuong-Nam
Nguyen-The, Quang
Vu-Minh, An
Nguyen, Diep-Anh
Pham, Xuan-Lam
Artificial Intelligence
History and Overview
I.2.7; K.3.1; G.3
ChatGPT, an Artificial Intelligence model, has the potential to revolutionize education. However, its effectiveness in solving non-English questions remains uncertain. This study evaluates ChatGPT's robustness using 586 Korean mathematics questions. ChatGPT achieves 66.72% accuracy, correctly answering 391 out of 586 questions. We also assess its ability to rate mathematics questions based on eleven criteria and perform a topic analysis. Our findings show that ChatGPT's ratings align with educational theory and test-taker perspectives. While ChatGPT performs well in question classification, it struggles with non-English contexts, highlighting areas for improvement. Future research should address linguistic biases and enhance accuracy across diverse languages. Domain-specific optimizations and multilingual training could improve ChatGPT's role in personalized education.
title On the robustness of ChatGPT in teaching Korean Mathematics
topic Artificial Intelligence
History and Overview
I.2.7; K.3.1; G.3
url https://arxiv.org/abs/2502.11915