Long Code Arena: a Set of Benchmarks for Long-Context Code Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bogomolov, Egor, Eliseeva, Aleksandra, Galimzyanov, Timur, Glukhov, Evgeniy, Shapkin, Anton, Tigina, Maria, Golubev, Yaroslav, Kovrigin, Alexander, van Deursen, Arie, Izadi, Maliheh, Bryksin, Timofey |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
by: Kovrigin, Alexander, et al.
Published: (2024)
by: Kovrigin, Alexander, et al.
Published: (2024)
Dynamic Retrieval-Augmented Generation
by: Shapkin, Anton, et al.
Published: (2023)
by: Shapkin, Anton, et al.
Published: (2023)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
by: Cipollone, Daniele, et al.
Published: (2025)
by: Cipollone, Daniele, et al.
Published: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024)
by: Galimzyanov, Timur, et al.
Published: (2024)
EnvBench: A Benchmark for Automated Environment Setup
by: Eliseeva, Aleksandra, et al.
Published: (2025)
by: Eliseeva, Aleksandra, et al.
Published: (2025)
PIPer: On-Device Environment Setup via Online Reinforcement Learning
by: Kovrigin, Alexander, et al.
Published: (2025)
by: Kovrigin, Alexander, et al.
Published: (2025)
Challenge on Optimization of Context Collection for Code Completion
by: Ustalov, Dmitry, et al.
Published: (2025)
by: Ustalov, Dmitry, et al.
Published: (2025)
Practical Code RAG at Scale: Task-Aware Retrieval Design Choices under Compute Budgets
by: Galimzyanov, Timur, et al.
Published: (2025)
by: Galimzyanov, Timur, et al.
Published: (2025)
Traces of Memorisation in Large Language Models for Code
by: Al-Kaswan, Ali, et al.
Published: (2023)
by: Al-Kaswan, Ali, et al.
Published: (2023)
Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings
by: Tsvetkov, Petr, et al.
Published: (2024)
by: Tsvetkov, Petr, et al.
Published: (2024)
A Transformer-Based Approach for Smart Invocation of Automatic Code Completion
by: de Moor, Aral, et al.
Published: (2024)
by: de Moor, Aral, et al.
Published: (2024)
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
by: Glukhov, Evgeniy, et al.
Published: (2025)
by: Glukhov, Evgeniy, et al.
Published: (2025)
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
by: Katzy, Jonathan, et al.
Published: (2024)
by: Katzy, Jonathan, et al.
Published: (2024)
The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
Using AI-Based Coding Assistants in Practice: State of Affairs, Perceptions, and Ways Forward
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
by: Al-Kaswan, Ali, et al.
Published: (2025)
by: Al-Kaswan, Ali, et al.
Published: (2025)
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
Clustering MOOC Programming Solutions to Diversify Their Presentation to Students
by: Artser, Elizaveta, et al.
Published: (2024)
by: Artser, Elizaveta, et al.
Published: (2024)
Developer Needs and Feasible Features for AI Assistants in IDEs
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Language Models for Code Completion: A Practical Evaluation
by: Izadi, Maliheh, et al.
Published: (2024)
by: Izadi, Maliheh, et al.
Published: (2024)
On Pretraining for Project-Level Code Completion
by: Sapronov, Maksim, et al.
Published: (2025)
by: Sapronov, Maksim, et al.
Published: (2025)
The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
by: Bagirov, Farid, et al.
Published: (2025)
by: Bagirov, Farid, et al.
Published: (2025)
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
Automated Attention Pattern Discovery at Scale in Large Language Models
by: Katzy, Jonathan, et al.
Published: (2026)
by: Katzy, Jonathan, et al.
Published: (2026)
Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
by: Mundhra, Yash, et al.
Published: (2025)
by: Mundhra, Yash, et al.
Published: (2025)
Together We Go Further: LLMs and IDE Static Analysis for Extract Method Refactoring
by: Pomian, Dorin, et al.
Published: (2024)
by: Pomian, Dorin, et al.
Published: (2024)
Kotlin ML Pack: Technical Report
by: Titov, Sergey, et al.
Published: (2024)
by: Titov, Sergey, et al.
Published: (2024)
Unprecedented Code Change Automation: The Fusion of LLMs and Transformation by Example
by: Dilhara, Malinda, et al.
Published: (2024)
by: Dilhara, Malinda, et al.
Published: (2024)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
Untangling Knots: Leveraging LLM for Error Resolution in Computational Notebooks
by: Grotov, Konstantin, et al.
Published: (2024)
by: Grotov, Konstantin, et al.
Published: (2024)
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
by: Lindenbauer, Tobias, et al.
Published: (2025)
by: Lindenbauer, Tobias, et al.
Published: (2025)
EM-Assist: Safe Automated ExtractMethod Refactoring with LLMs
by: Pomian, Dorin, et al.
Published: (2024)
by: Pomian, Dorin, et al.
Published: (2024)
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
by: Lindenbauer, Tobias, et al.
Published: (2025)
by: Lindenbauer, Tobias, et al.
Published: (2025)
Assessing Consensus of Developers' Views on Code Readability
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
AST-PAC: AST-guided Membership Inference for Code
by: Koohestani, Roham, et al.
Published: (2026)
by: Koohestani, Roham, et al.
Published: (2026)
Debug Smarter, Not Harder: AI Agents for Error Resolution in Computational Notebooks
by: Grotov, Konstantin, et al.
Published: (2024)
by: Grotov, Konstantin, et al.
Published: (2024)
How Much Do Code Language Models Remember? An Investigation on Data Extraction Attacks before and after Fine-tuning
by: Salerno, Fabio, et al.
Published: (2025)
by: Salerno, Fabio, et al.
Published: (2025)
Context Composing for Full Line Code Completion
by: Semenkin, Anton, et al.
Published: (2024)
by: Semenkin, Anton, et al.
Published: (2024)
From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution
by: Chizhov, Pavel, et al.
Published: (2026)
by: Chizhov, Pavel, et al.
Published: (2026)
Full Line Code Completion: Bringing AI to Desktop
by: Semenkin, Anton, et al.
Published: (2024)
by: Semenkin, Anton, et al.
Published: (2024)
Similar Items
-
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
by: Kovrigin, Alexander, et al.
Published: (2024) -
Dynamic Retrieval-Augmented Generation
by: Shapkin, Anton, et al.
Published: (2023) -
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
by: Cipollone, Daniele, et al.
Published: (2025) -
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024) -
EnvBench: A Benchmark for Automated Environment Setup
by: Eliseeva, Aleksandra, et al.
Published: (2025)