Long Code Arena: a Set of Benchmarks for Long-Context Code Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bogomolov, Egor, Eliseeva, Aleksandra, Galimzyanov, Timur, Glukhov, Evgeniy, Shapkin, Anton, Tigina, Maria, Golubev, Yaroslav, Kovrigin, Alexander, van Deursen, Arie, Izadi, Maliheh, Bryksin, Timofey
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911920738533376
author Bogomolov, Egor
Eliseeva, Aleksandra
Galimzyanov, Timur
Glukhov, Evgeniy
Shapkin, Anton
Tigina, Maria
Golubev, Yaroslav
Kovrigin, Alexander
van Deursen, Arie
Izadi, Maliheh
Bryksin, Timofey
author_facet Bogomolov, Egor
Eliseeva, Aleksandra
Galimzyanov, Timur
Glukhov, Evgeniy
Shapkin, Anton
Tigina, Maria
Golubev, Yaroslav
Kovrigin, Alexander
van Deursen, Arie
Izadi, Maliheh
Bryksin, Timofey
contents Nowadays, the fields of code and natural language processing are evolving rapidly. In particular, models become better at processing long context windows - supported context sizes have increased by orders of magnitude over the last few years. However, there is a shortage of benchmarks for code processing that go beyond a single file of context, while the most popular ones are limited to a single method. With this work, we aim to close this gap by introducing Long Code Arena, a suite of six benchmarks for code processing tasks that require project-wide context. These tasks cover different aspects of code processing: library-based code generation, CI builds repair, project-level code completion, commit message generation, bug localization, and module summarization. For each task, we provide a manually verified dataset for testing, an evaluation suite, and open-source baseline solutions based on popular LLMs to showcase the usage of the dataset and to simplify adoption by other researchers. We publish the benchmark page on HuggingFace Spaces with the leaderboard, links to HuggingFace Hub for all the datasets, and link to the GitHub repository with baselines: https://huggingface.co/spaces/JetBrains-Research/long-code-arena.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11612
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Long Code Arena: a Set of Benchmarks for Long-Context Code Models
Bogomolov, Egor
Eliseeva, Aleksandra
Galimzyanov, Timur
Glukhov, Evgeniy
Shapkin, Anton
Tigina, Maria
Golubev, Yaroslav
Kovrigin, Alexander
van Deursen, Arie
Izadi, Maliheh
Bryksin, Timofey
Machine Learning
Artificial Intelligence
Information Retrieval
Software Engineering
Nowadays, the fields of code and natural language processing are evolving rapidly. In particular, models become better at processing long context windows - supported context sizes have increased by orders of magnitude over the last few years. However, there is a shortage of benchmarks for code processing that go beyond a single file of context, while the most popular ones are limited to a single method. With this work, we aim to close this gap by introducing Long Code Arena, a suite of six benchmarks for code processing tasks that require project-wide context. These tasks cover different aspects of code processing: library-based code generation, CI builds repair, project-level code completion, commit message generation, bug localization, and module summarization. For each task, we provide a manually verified dataset for testing, an evaluation suite, and open-source baseline solutions based on popular LLMs to showcase the usage of the dataset and to simplify adoption by other researchers. We publish the benchmark page on HuggingFace Spaces with the leaderboard, links to HuggingFace Hub for all the datasets, and link to the GitHub repository with baselines: https://huggingface.co/spaces/JetBrains-Research/long-code-arena.
title Long Code Arena: a Set of Benchmarks for Long-Context Code Models
topic Machine Learning
Artificial Intelligence
Information Retrieval
Software Engineering
url https://arxiv.org/abs/2406.11612