Saved in:
| Main Authors: | Xiao, Yijia, Wang, Runhui, Kong, Luyang, Golac, Davor, Wang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.06111 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization
by: Wang, Juntong, et al.
Published: (2026)
by: Wang, Juntong, et al.
Published: (2026)
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
by: Wang, Lilin, et al.
Published: (2025)
by: Wang, Lilin, et al.
Published: (2025)
LLM-based Content Classification Approach for GitHub Repositories by the README Files
by: Mehmood, Malik Uzair, et al.
Published: (2025)
by: Mehmood, Malik Uzair, et al.
Published: (2025)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
by: Rank, Ben, et al.
Published: (2026)
by: Rank, Ben, et al.
Published: (2026)
AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems
by: Wang, Zhaohui Geoffrey
Published: (2026)
by: Wang, Zhaohui Geoffrey
Published: (2026)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
by: Duston, Titouan, et al.
Published: (2025)
by: Duston, Titouan, et al.
Published: (2025)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
by: Xu, WeiZhe, et al.
Published: (2026)
by: Xu, WeiZhe, et al.
Published: (2026)
SoundnessBench: A Soundness Benchmark for Neural Network Verifiers
by: Zhou, Xingjian, et al.
Published: (2024)
by: Zhou, Xingjian, et al.
Published: (2024)
SWE-Bench-CL: Continual Learning for Coding Agents
by: Joshi, Thomas, et al.
Published: (2025)
by: Joshi, Thomas, et al.
Published: (2025)
AIPC: Agent-Based Automation for AI Model Deployment with Qualcomm AI Runtime
by: Su, Jianhao, et al.
Published: (2026)
by: Su, Jianhao, et al.
Published: (2026)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
by: Shbita, Basel, et al.
Published: (2025)
by: Shbita, Basel, et al.
Published: (2025)
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
by: Arora, Avi, et al.
Published: (2025)
by: Arora, Avi, et al.
Published: (2025)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
by: Mündler, Niels, et al.
Published: (2024)
by: Mündler, Niels, et al.
Published: (2024)
DRAGON: Robust Classification for Very Large Collections of Software Repositories
by: Balla, Stefano, et al.
Published: (2026)
by: Balla, Stefano, et al.
Published: (2026)
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
by: Kovrigin, Alexander, et al.
Published: (2024)
by: Kovrigin, Alexander, et al.
Published: (2024)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2025)
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2025)
QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture
by: Prakash, Shvetank, et al.
Published: (2025)
by: Prakash, Shvetank, et al.
Published: (2025)
The BrowserGym Ecosystem for Web Agent Research
by: De Chezelles, Thibault Le Sellier, et al.
Published: (2024)
by: De Chezelles, Thibault Le Sellier, et al.
Published: (2024)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
GRAM: Generative Retrieval Augmented Matching of Data Schemas in the Context of Data Security
by: Liu, Xuanqing, et al.
Published: (2024)
by: Liu, Xuanqing, et al.
Published: (2024)
The Dual-State Architecture for Reliable LLM Agents
by: Thompson, Matthew
Published: (2025)
by: Thompson, Matthew
Published: (2025)
DafnyBench: A Benchmark for Formal Software Verification
by: Loughridge, Chloe, et al.
Published: (2024)
by: Loughridge, Chloe, et al.
Published: (2024)
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
by: Xie, Zichen, et al.
Published: (2026)
by: Xie, Zichen, et al.
Published: (2026)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
by: Wang, Yubang, et al.
Published: (2026)
by: Wang, Yubang, et al.
Published: (2026)
Repo2Run: Automated Building Executable Environment for Code Repository at Scale
by: Hu, Ruida, et al.
Published: (2025)
by: Hu, Ruida, et al.
Published: (2025)
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
by: Yu, Shasha, et al.
Published: (2026)
by: Yu, Shasha, et al.
Published: (2026)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
by: Qiu, Ruizhong, et al.
Published: (2024)
by: Qiu, Ruizhong, et al.
Published: (2024)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
by: Zheng, Qinkai, et al.
Published: (2023)
by: Zheng, Qinkai, et al.
Published: (2023)
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
by: Feng, Yunfei, et al.
Published: (2026)
by: Feng, Yunfei, et al.
Published: (2026)
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
by: Jha, Smriti, et al.
Published: (2026)
by: Jha, Smriti, et al.
Published: (2026)
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
by: Manglik, Akshay, et al.
Published: (2026)
by: Manglik, Akshay, et al.
Published: (2026)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
by: Vulićević, Jelena Ilić
Published: (2026)
by: Vulićević, Jelena Ilić
Published: (2026)
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
by: Han, Xiaoke, et al.
Published: (2025)
by: Han, Xiaoke, et al.
Published: (2025)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
Deploying Geospatial Foundation Models in the Real World: Lessons from WorldCereal
by: Butsko, Christina, et al.
Published: (2025)
by: Butsko, Christina, et al.
Published: (2025)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
by: Yuan, Danlong, et al.
Published: (2026)
by: Yuan, Danlong, et al.
Published: (2026)
Relative Positioning Based Code Chunking Method For Rich Context Retrieval In Repository Level Code Completion Task With Code Language Model
by: Rahman, Imranur, et al.
Published: (2025)
by: Rahman, Imranur, et al.
Published: (2025)
Similar Items
-
GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization
by: Wang, Juntong, et al.
Published: (2026) -
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
by: Wang, Lilin, et al.
Published: (2025) -
LLM-based Content Classification Approach for GitHub Repositories by the README Files
by: Mehmood, Malik Uzair, et al.
Published: (2025) -
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
by: Rank, Ben, et al.
Published: (2026) -
AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems
by: Wang, Zhaohui Geoffrey
Published: (2026)