YABLoCo: Yet Another Benchmark for Long Context Code Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Valeev, Aidar, Garaev, Roman, Lomshakov, Vadim, Piontkovskaya, Irina, Ivanov, Vladimir, Adewuyi, Israel
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915276253036544
author Valeev, Aidar
Garaev, Roman
Lomshakov, Vadim
Piontkovskaya, Irina
Ivanov, Vladimir
Adewuyi, Israel
author_facet Valeev, Aidar
Garaev, Roman
Lomshakov, Vadim
Piontkovskaya, Irina
Ivanov, Vladimir
Adewuyi, Israel
contents Large Language Models demonstrate the ability to solve various programming tasks, including code generation. Typically, the performance of LLMs is measured on benchmarks with small or medium-sized context windows of thousands of lines of code. At the same time, in real-world software projects, repositories can span up to millions of LoC. This paper closes this gap by contributing to the long context code generation benchmark (YABLoCo). The benchmark featured a test set of 215 functions selected from four large repositories with thousands of functions. The dataset contained metadata of functions, contexts of the functions with different levels of dependencies, docstrings, functions bodies, and call graphs for each repository. This paper presents three key aspects of the contribution. First, the benchmark aims at function body generation in large repositories in C and C++, two languages not covered by previous benchmarks. Second, the benchmark contains large repositories from 200K to 2,000K LoC. Third, we contribute a scalable evaluation pipeline for efficient computing of the target metrics and a tool for visual analysis of generated code. Overall, these three aspects allow for evaluating code generation in large repositories in C and C++.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle YABLoCo: Yet Another Benchmark for Long Context Code Generation
Valeev, Aidar
Garaev, Roman
Lomshakov, Vadim
Piontkovskaya, Irina
Ivanov, Vladimir
Adewuyi, Israel
Computation and Language
Artificial Intelligence
Software Engineering
Large Language Models demonstrate the ability to solve various programming tasks, including code generation. Typically, the performance of LLMs is measured on benchmarks with small or medium-sized context windows of thousands of lines of code. At the same time, in real-world software projects, repositories can span up to millions of LoC. This paper closes this gap by contributing to the long context code generation benchmark (YABLoCo). The benchmark featured a test set of 215 functions selected from four large repositories with thousands of functions. The dataset contained metadata of functions, contexts of the functions with different levels of dependencies, docstrings, functions bodies, and call graphs for each repository. This paper presents three key aspects of the contribution. First, the benchmark aims at function body generation in large repositories in C and C++, two languages not covered by previous benchmarks. Second, the benchmark contains large repositories from 200K to 2,000K LoC. Third, we contribute a scalable evaluation pipeline for efficient computing of the target metrics and a tool for visual analysis of generated code. Overall, these three aspects allow for evaluating code generation in large repositories in C and C++.
title YABLoCo: Yet Another Benchmark for Long Context Code Generation
topic Computation and Language
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2505.04406