MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chervyakov, Artem, Kharitonov, Alexander, Zadorozhny, Pavel, Pavel, Adamenko, Levichev, Rodion, Vorobev, Dmitrii, Salikhov, Dmitrii, Valeev, Aidar, Pestova, Alena, Dziuba, Maria, Alimova, Ilseyar, Zavgorodnev, Artem, Medvedev, Aleksandr, Moiseev, Stanislav, Bruches, Elena, Grebenkin, Daniil, Derunets, Roman, Vladimir, Vikulov, Emelyanov, Anton, Babaev, Dmitrii, Ivanov, Vladimir V., Malykh, Valentin, Fenogenova, Alena
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909936113418240
author Chervyakov, Artem
Kharitonov, Alexander
Zadorozhny, Pavel
Pavel, Adamenko
Levichev, Rodion
Vorobev, Dmitrii
Salikhov, Dmitrii
Valeev, Aidar
Pestova, Alena
Dziuba, Maria
Alimova, Ilseyar
Zavgorodnev, Artem
Medvedev, Aleksandr
Moiseev, Stanislav
Bruches, Elena
Grebenkin, Daniil
Derunets, Roman
Vladimir, Vikulov
Emelyanov, Anton
Babaev, Dmitrii
Ivanov, Vladimir V.
Malykh, Valentin
Fenogenova, Alena
author_facet Chervyakov, Artem
Kharitonov, Alexander
Zadorozhny, Pavel
Pavel, Adamenko
Levichev, Rodion
Vorobev, Dmitrii
Salikhov, Dmitrii
Valeev, Aidar
Pestova, Alena
Dziuba, Maria
Alimova, Ilseyar
Zavgorodnev, Artem
Medvedev, Aleksandr
Moiseev, Stanislav
Bruches, Elena
Grebenkin, Daniil
Derunets, Roman
Vladimir, Vikulov
Emelyanov, Anton
Babaev, Dmitrii
Ivanov, Vladimir V.
Malykh, Valentin
Fenogenova, Alena
contents Advancements in LLMs have enhanced task automation in software engineering; however, current evaluations primarily focus on natural language tasks, overlooking code quality. Most benchmarks prioritize high-level reasoning over executable code and real-world performance, leaving gaps in understanding true capabilities and risks associated with these models in production. To address this issue, we propose MERA Code, a new addition to the MERA benchmark family, specifically focused on evaluating code for the latest code generation LLMs in Russian. This benchmark includes 11 evaluation tasks that span 8 programming languages. Our proposed evaluation methodology features a taxonomy that outlines the practical coding skills necessary for models to complete these tasks. The benchmark comprises an open-source codebase for users to conduct MERA assessments, a scoring system compatible with various programming environments, and a platform featuring a leaderboard and submission system. We evaluate open LLMs and frontier API models, analyzing their limitations in terms of practical coding tasks in non-English languages. We are publicly releasing MERA to guide future research, anticipate groundbreaking features in model development, and standardize evaluation procedures.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12284
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks
Chervyakov, Artem
Kharitonov, Alexander
Zadorozhny, Pavel
Pavel, Adamenko
Levichev, Rodion
Vorobev, Dmitrii
Salikhov, Dmitrii
Valeev, Aidar
Pestova, Alena
Dziuba, Maria
Alimova, Ilseyar
Zavgorodnev, Artem
Medvedev, Aleksandr
Moiseev, Stanislav
Bruches, Elena
Grebenkin, Daniil
Derunets, Roman
Vladimir, Vikulov
Emelyanov, Anton
Babaev, Dmitrii
Ivanov, Vladimir V.
Malykh, Valentin
Fenogenova, Alena
Software Engineering
Artificial Intelligence
Computation and Language
Advancements in LLMs have enhanced task automation in software engineering; however, current evaluations primarily focus on natural language tasks, overlooking code quality. Most benchmarks prioritize high-level reasoning over executable code and real-world performance, leaving gaps in understanding true capabilities and risks associated with these models in production. To address this issue, we propose MERA Code, a new addition to the MERA benchmark family, specifically focused on evaluating code for the latest code generation LLMs in Russian. This benchmark includes 11 evaluation tasks that span 8 programming languages. Our proposed evaluation methodology features a taxonomy that outlines the practical coding skills necessary for models to complete these tasks. The benchmark comprises an open-source codebase for users to conduct MERA assessments, a scoring system compatible with various programming environments, and a platform featuring a leaderboard and submission system. We evaluate open LLMs and frontier API models, analyzing their limitations in terms of practical coding tasks in non-English languages. We are publicly releasing MERA to guide future research, anticipate groundbreaking features in model development, and standardize evaluation procedures.
title MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.12284