Evaluating Non-English Developer Support in Machine Learning for Software Engineering

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Katzy, Jonathan, Huang, Yongcheng, Panchu, Gopal-Raj, Ziemlewski, Maksym, Loizides, Paris, Vermeulen, Sander, van Deursen, Arie, Izadi, Maliheh
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914537772417024
author Katzy, Jonathan
Huang, Yongcheng
Panchu, Gopal-Raj
Ziemlewski, Maksym
Loizides, Paris
Vermeulen, Sander
van Deursen, Arie
Izadi, Maliheh
author_facet Katzy, Jonathan
Huang, Yongcheng
Panchu, Gopal-Raj
Ziemlewski, Maksym
Loizides, Paris
Vermeulen, Sander
van Deursen, Arie
Izadi, Maliheh
contents Large Language Models are increasingly used in software engineering, but both code generation and its evaluation remain predominantly English-centric. This leaves a major gap in our understanding of how well current tools support multilingual development, where code contains non-English natural language. In this paper, we investigate non-English code comment generation and the reliability of current methods for evaluating such outputs. We evaluate five code LLMs (CodeGemma, CodeLlama, CodeQwen1.5, GraniteCode, and StarCoder2) across five natural languages: Dutch, English, Greek, Polish and Chinese. We further conduct an open-coding study of 12,500 generated comments, from which we derive a publicly released human-annotated dataset and a taxonomy of 26 error types. We use these human annotations, to evaluate the performance of neural metrics, and LLM-as-a-judge pipelines. Our findings show that generative performance deteriorates substantially outside English, with linguistic errors increasing by up to 15.1$\times$, alongside frequent incoherent generations and a rise in semantic errors. More critically, we show that detecting errors in non-English comments underperforms. Across classical overlap-based metrics, off-the-shelf neural metrics, extended neural metrics using newer multilingual, language-specific, and code-specific models, and LLM-as-a-judge pipelines, no automatic approach provides reliable and consistent assessment. Neural metrics fail to distinguish correct comments from incorrect outputs or even random noise, and tend to overestimate quality in non-English settings. LLM-as-a-judge methods achieve the highest agreement with human annotations but fail to reliably capture important language-related and semantic errors. Overall, our results show that evaluation and generation are key barriers for multilingual tooling, and that human judgment remains indispensable.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05902
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating Non-English Developer Support in Machine Learning for Software Engineering
Katzy, Jonathan
Huang, Yongcheng
Panchu, Gopal-Raj
Ziemlewski, Maksym
Loizides, Paris
Vermeulen, Sander
van Deursen, Arie
Izadi, Maliheh
Software Engineering
Large Language Models are increasingly used in software engineering, but both code generation and its evaluation remain predominantly English-centric. This leaves a major gap in our understanding of how well current tools support multilingual development, where code contains non-English natural language. In this paper, we investigate non-English code comment generation and the reliability of current methods for evaluating such outputs. We evaluate five code LLMs (CodeGemma, CodeLlama, CodeQwen1.5, GraniteCode, and StarCoder2) across five natural languages: Dutch, English, Greek, Polish and Chinese. We further conduct an open-coding study of 12,500 generated comments, from which we derive a publicly released human-annotated dataset and a taxonomy of 26 error types. We use these human annotations, to evaluate the performance of neural metrics, and LLM-as-a-judge pipelines. Our findings show that generative performance deteriorates substantially outside English, with linguistic errors increasing by up to 15.1$\times$, alongside frequent incoherent generations and a rise in semantic errors. More critically, we show that detecting errors in non-English comments underperforms. Across classical overlap-based metrics, off-the-shelf neural metrics, extended neural metrics using newer multilingual, language-specific, and code-specific models, and LLM-as-a-judge pipelines, no automatic approach provides reliable and consistent assessment. Neural metrics fail to distinguish correct comments from incorrect outputs or even random noise, and tend to overestimate quality in non-English settings. LLM-as-a-judge methods achieve the highest agreement with human annotations but fail to reliably capture important language-related and semantic errors. Overall, our results show that evaluation and generation are key barriers for multilingual tooling, and that human judgment remains indispensable.
title Evaluating Non-English Developer Support in Machine Learning for Software Engineering
topic Software Engineering
url https://arxiv.org/abs/2605.05902