EnvBench: A Benchmark for Automated Environment Setup
Fuente:
arXiv
Saved in:
| Main Authors: | Eliseeva, Aleksandra, Kovrigin, Alexander, Kholkin, Ilia, Bogomolov, Egor, Zharov, Yaroslav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PIPer: On-Device Environment Setup via Online Reinforcement Learning
by: Kovrigin, Alexander, et al.
Published: (2025)
by: Kovrigin, Alexander, et al.
Published: (2025)
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
by: Kovrigin, Alexander, et al.
Published: (2024)
by: Kovrigin, Alexander, et al.
Published: (2024)
Step Rejection Fine-Tuning: A Practical Distillation Recipe
by: Slinko, Igor, et al.
Published: (2026)
by: Slinko, Igor, et al.
Published: (2026)
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
by: Lindenbauer, Tobias, et al.
Published: (2025)
by: Lindenbauer, Tobias, et al.
Published: (2025)
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
by: Bogomolov, Egor, et al.
Published: (2024)
by: Bogomolov, Egor, et al.
Published: (2024)
Dynamic Retrieval-Augmented Generation
by: Shapkin, Anton, et al.
Published: (2023)
by: Shapkin, Anton, et al.
Published: (2023)
Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings
by: Tsvetkov, Petr, et al.
Published: (2024)
by: Tsvetkov, Petr, et al.
Published: (2024)
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
by: Glukhov, Evgeniy, et al.
Published: (2025)
by: Glukhov, Evgeniy, et al.
Published: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024)
by: Galimzyanov, Timur, et al.
Published: (2024)
On Problems of Implicit Context Compression for Software Engineering Agents
by: Gelvan, Kirill, et al.
Published: (2026)
by: Gelvan, Kirill, et al.
Published: (2026)
Tool-Augmented LLMs as a Universal Interface for IDEs
by: Zharov, Yaroslav, et al.
Published: (2024)
by: Zharov, Yaroslav, et al.
Published: (2024)
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
by: Lindenbauer, Tobias, et al.
Published: (2025)
by: Lindenbauer, Tobias, et al.
Published: (2025)
Untangling Knots: Leveraging LLM for Error Resolution in Computational Notebooks
by: Grotov, Konstantin, et al.
Published: (2024)
by: Grotov, Konstantin, et al.
Published: (2024)
Challenge on Optimization of Context Collection for Code Completion
by: Ustalov, Dmitry, et al.
Published: (2025)
by: Ustalov, Dmitry, et al.
Published: (2025)
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
by: Arora, Avi, et al.
Published: (2025)
by: Arora, Avi, et al.
Published: (2025)
Stack Trace Deduplication: Faster, More Accurately, and in More Realistic Scenarios
by: Shibaev, Egor, et al.
Published: (2024)
by: Shibaev, Egor, et al.
Published: (2024)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
by: Wang, Yubang, et al.
Published: (2026)
by: Wang, Yubang, et al.
Published: (2026)
EM-Assist: Safe Automated ExtractMethod Refactoring with LLMs
by: Pomian, Dorin, et al.
Published: (2024)
by: Pomian, Dorin, et al.
Published: (2024)
OSS-Bench: Benchmark Generator for Coding LLMs
by: Jiang, Yuancheng, et al.
Published: (2025)
by: Jiang, Yuancheng, et al.
Published: (2025)
Automating Benchmark Design
by: Dsouza, Amanda, et al.
Published: (2025)
by: Dsouza, Amanda, et al.
Published: (2025)
ThrowBench: Benchmarking LLMs by Predicting Runtime Exceptions
by: Prenner, Julian Aron, et al.
Published: (2025)
by: Prenner, Julian Aron, et al.
Published: (2025)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2026)
by: Orel, Daniil, et al.
Published: (2026)
AssertionBench: A Benchmark to Evaluate Large-Language Models for Assertion Generation
by: Pulavarthi, Vaishnavi, et al.
Published: (2024)
by: Pulavarthi, Vaishnavi, et al.
Published: (2024)
MobileDev-Bench: A Benchmark for Issue Resolution in Mobile Application Development
by: Fakorede, Moshood A., et al.
Published: (2026)
by: Fakorede, Moshood A., et al.
Published: (2026)
Context Composing for Full Line Code Completion
by: Semenkin, Anton, et al.
Published: (2024)
by: Semenkin, Anton, et al.
Published: (2024)
SoundnessBench: A Soundness Benchmark for Neural Network Verifiers
by: Zhou, Xingjian, et al.
Published: (2024)
by: Zhou, Xingjian, et al.
Published: (2024)
LLM-Enhanced Log Anomaly Detection: A Comprehensive Benchmark of Large Language Models for Automated System Diagnostics
by: Patel, Disha
Published: (2026)
by: Patel, Disha
Published: (2026)
CRUST-Bench: A Comprehensive Benchmark for C-to-safe-Rust Transpilation
by: Khatry, Anirudh, et al.
Published: (2025)
by: Khatry, Anirudh, et al.
Published: (2025)
Risk Management for Mitigating Benchmark Failure Modes: BenchRisk
by: McGregor, Sean, et al.
Published: (2025)
by: McGregor, Sean, et al.
Published: (2025)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
by: Rank, Ben, et al.
Published: (2026)
by: Rank, Ben, et al.
Published: (2026)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
by: Shbita, Basel, et al.
Published: (2025)
by: Shbita, Basel, et al.
Published: (2025)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
by: Xiao, Yijia, et al.
Published: (2025)
by: Xiao, Yijia, et al.
Published: (2025)
CoDocBench: A Dataset for Code-Documentation Alignment in Software Maintenance
by: Pai, Kunal, et al.
Published: (2025)
by: Pai, Kunal, et al.
Published: (2025)
RepairBench: Leaderboard of Frontier Models for Program Repair
by: Silva, André, et al.
Published: (2024)
by: Silva, André, et al.
Published: (2024)
TeleResilienceBench: Quantifying Resilience for LLM Reasoning in Telecommunications
by: Gajjar, Pranshav, et al.
Published: (2026)
by: Gajjar, Pranshav, et al.
Published: (2026)
DafnyBench: A Benchmark for Formal Software Verification
by: Loughridge, Chloe, et al.
Published: (2024)
by: Loughridge, Chloe, et al.
Published: (2024)
Full Line Code Completion: Bringing AI to Desktop
by: Semenkin, Anton, et al.
Published: (2024)
by: Semenkin, Anton, et al.
Published: (2024)
Similar Items
-
PIPer: On-Device Environment Setup via Online Reinforcement Learning
by: Kovrigin, Alexander, et al.
Published: (2025) -
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
by: Kovrigin, Alexander, et al.
Published: (2024) -
Step Rejection Fine-Tuning: A Practical Distillation Recipe
by: Slinko, Igor, et al.
Published: (2026) -
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
by: Lindenbauer, Tobias, et al.
Published: (2025) -
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
by: Bogomolov, Egor, et al.
Published: (2024)