CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Yijia, Wang, Runhui, Kong, Luyang, Golac, Davor, Wang, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization
von: Wang, Juntong, et al.
Veröffentlicht: (2026)
von: Wang, Juntong, et al.
Veröffentlicht: (2026)
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
von: Wang, Lilin, et al.
Veröffentlicht: (2025)
von: Wang, Lilin, et al.
Veröffentlicht: (2025)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
von: Rank, Ben, et al.
Veröffentlicht: (2026)
von: Rank, Ben, et al.
Veröffentlicht: (2026)
LLM-based Content Classification Approach for GitHub Repositories by the README Files
von: Mehmood, Malik Uzair, et al.
Veröffentlicht: (2025)
von: Mehmood, Malik Uzair, et al.
Veröffentlicht: (2025)
AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems
von: Wang, Zhaohui Geoffrey
Veröffentlicht: (2026)
von: Wang, Zhaohui Geoffrey
Veröffentlicht: (2026)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025)
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
SoundnessBench: A Soundness Benchmark for Neural Network Verifiers
von: Zhou, Xingjian, et al.
Veröffentlicht: (2024)
von: Zhou, Xingjian, et al.
Veröffentlicht: (2024)
SWE-Bench-CL: Continual Learning for Coding Agents
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
von: Xu, WeiZhe, et al.
Veröffentlicht: (2026)
von: Xu, WeiZhe, et al.
Veröffentlicht: (2026)
AIPC: Agent-Based Automation for AI Model Deployment with Qualcomm AI Runtime
von: Su, Jianhao, et al.
Veröffentlicht: (2026)
von: Su, Jianhao, et al.
Veröffentlicht: (2026)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
von: Shbita, Basel, et al.
Veröffentlicht: (2025)
von: Shbita, Basel, et al.
Veröffentlicht: (2025)
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
von: Arora, Avi, et al.
Veröffentlicht: (2025)
von: Arora, Avi, et al.
Veröffentlicht: (2025)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
DRAGON: Robust Classification for Very Large Collections of Software Repositories
von: Balla, Stefano, et al.
Veröffentlicht: (2026)
von: Balla, Stefano, et al.
Veröffentlicht: (2026)
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
von: Kovrigin, Alexander, et al.
Veröffentlicht: (2024)
von: Kovrigin, Alexander, et al.
Veröffentlicht: (2024)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
The BrowserGym Ecosystem for Web Agent Research
von: De Chezelles, Thibault Le Sellier, et al.
Veröffentlicht: (2024)
von: De Chezelles, Thibault Le Sellier, et al.
Veröffentlicht: (2024)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2025)
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2025)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
von: Wang, Yubang, et al.
Veröffentlicht: (2026)
von: Wang, Yubang, et al.
Veröffentlicht: (2026)
QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture
von: Prakash, Shvetank, et al.
Veröffentlicht: (2025)
von: Prakash, Shvetank, et al.
Veröffentlicht: (2025)
The Dual-State Architecture for Reliable LLM Agents
von: Thompson, Matthew
Veröffentlicht: (2025)
von: Thompson, Matthew
Veröffentlicht: (2025)
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
von: Yu, Shasha, et al.
Veröffentlicht: (2026)
von: Yu, Shasha, et al.
Veröffentlicht: (2026)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
von: Qiu, Ruizhong, et al.
Veröffentlicht: (2024)
von: Qiu, Ruizhong, et al.
Veröffentlicht: (2024)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
von: Feng, Yunfei, et al.
Veröffentlicht: (2026)
von: Feng, Yunfei, et al.
Veröffentlicht: (2026)
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
von: Jha, Smriti, et al.
Veröffentlicht: (2026)
von: Jha, Smriti, et al.
Veröffentlicht: (2026)
DafnyBench: A Benchmark for Formal Software Verification
von: Loughridge, Chloe, et al.
Veröffentlicht: (2024)
von: Loughridge, Chloe, et al.
Veröffentlicht: (2024)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
von: Manglik, Akshay, et al.
Veröffentlicht: (2026)
von: Manglik, Akshay, et al.
Veröffentlicht: (2026)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
von: Han, Xiaoke, et al.
Veröffentlicht: (2025)
von: Han, Xiaoke, et al.
Veröffentlicht: (2025)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
von: Rahman, Musfiqur, et al.
Veröffentlicht: (2025)
von: Rahman, Musfiqur, et al.
Veröffentlicht: (2025)
Repo2Run: Automated Building Executable Environment for Code Repository at Scale
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
Deploying Geospatial Foundation Models in the Real World: Lessons from WorldCereal
von: Butsko, Christina, et al.
Veröffentlicht: (2025)
von: Butsko, Christina, et al.
Veröffentlicht: (2025)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
von: Singh, Jaskirat, et al.
Veröffentlicht: (2024)
von: Singh, Jaskirat, et al.
Veröffentlicht: (2024)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
von: Diehl, Patrick, et al.
Veröffentlicht: (2025)
von: Diehl, Patrick, et al.
Veröffentlicht: (2025)
Relative Positioning Based Code Chunking Method For Rich Context Retrieval In Repository Level Code Completion Task With Code Language Model
von: Rahman, Imranur, et al.
Veröffentlicht: (2025)
von: Rahman, Imranur, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization
von: Wang, Juntong, et al.
Veröffentlicht: (2026) -
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
von: Wang, Lilin, et al.
Veröffentlicht: (2025) -
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
von: Rank, Ben, et al.
Veröffentlicht: (2026) -
LLM-based Content Classification Approach for GitHub Repositories by the README Files
von: Mehmood, Malik Uzair, et al.
Veröffentlicht: (2025) -
AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems
von: Wang, Zhaohui Geoffrey
Veröffentlicht: (2026)