SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Petrukha, Ivan, Kurliak, Yana, Stulova, Nataliia
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913866938580992
author Petrukha, Ivan
Kurliak, Yana
Stulova, Nataliia
author_facet Petrukha, Ivan
Kurliak, Yana
Stulova, Nataliia
contents In recent years, large language models (LLMs) have showcased significant advancements in code generation. However, most evaluation benchmarks are primarily oriented towards Python, making it difficult to evaluate other programming languages, such as Swift, with high quality. By examining widely established multilingual benchmarks like HumanEval-XL and MultiPL-E, we identified critical issues specific to their Swift components, making them insufficient or even irrelevant for assessing LLM coding capabilities on Swift. Unlike these existing approaches, which prioritize rapid scaling and generalization by automatically translating Python-centric benchmarks with LLMs, we adopt a quality-over-quantity methodology. We present SwiftEval, the first Swift-oriented benchmark consisting of 28 carefully hand-crafted problems, and evaluate 44 popular Code LLMs on it. Our results show significant LLM scores drop for problems requiring language-specific features, most noticeable in the models of smaller sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24324
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
Petrukha, Ivan
Kurliak, Yana
Stulova, Nataliia
Machine Learning
Computation and Language
Programming Languages
Software Engineering
In recent years, large language models (LLMs) have showcased significant advancements in code generation. However, most evaluation benchmarks are primarily oriented towards Python, making it difficult to evaluate other programming languages, such as Swift, with high quality. By examining widely established multilingual benchmarks like HumanEval-XL and MultiPL-E, we identified critical issues specific to their Swift components, making them insufficient or even irrelevant for assessing LLM coding capabilities on Swift. Unlike these existing approaches, which prioritize rapid scaling and generalization by automatically translating Python-centric benchmarks with LLMs, we adopt a quality-over-quantity methodology. We present SwiftEval, the first Swift-oriented benchmark consisting of 28 carefully hand-crafted problems, and evaluate 44 popular Code LLMs on it. Our results show significant LLM scores drop for problems requiring language-specific features, most noticeable in the models of smaller sizes.
title SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
topic Machine Learning
Computation and Language
Programming Languages
Software Engineering
url https://arxiv.org/abs/2505.24324