UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Boxi, Zhu, Yuxuan, He, Pinjia, Kang, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey
von: Li, Caihua, et al.
Veröffentlicht: (2026)
von: Li, Caihua, et al.
Veröffentlicht: (2026)
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
von: Li, Kenan, et al.
Veröffentlicht: (2026)
von: Li, Kenan, et al.
Veröffentlicht: (2026)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
von: Dinu, Ion George, et al.
Veröffentlicht: (2026)
von: Dinu, Ion George, et al.
Veröffentlicht: (2026)
Automated Deep Learning Optimization via DSL-Based Source Code Transformation
von: Wang, Ruixin, et al.
Veröffentlicht: (2024)
von: Wang, Ruixin, et al.
Veröffentlicht: (2024)
CIFE: Code Instruction-Following Evaluation
von: Gunnu, Sravani, et al.
Veröffentlicht: (2025)
von: Gunnu, Sravani, et al.
Veröffentlicht: (2025)
AutoCodeSherpa: Symbolic Explanations in AI Coding Agents
von: Kang, Sungmin, et al.
Veröffentlicht: (2025)
von: Kang, Sungmin, et al.
Veröffentlicht: (2025)
The Right Prompts for the Job: Repair Code-Review Defects with Large Language Model
von: Zhao, Zelin, et al.
Veröffentlicht: (2023)
von: Zhao, Zelin, et al.
Veröffentlicht: (2023)
RepoAudit: An Autonomous LLM-Agent for Repository-Level Code Auditing
von: Guo, Jinyao, et al.
Veröffentlicht: (2025)
von: Guo, Jinyao, et al.
Veröffentlicht: (2025)
Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures
von: Rombaut, Benjamin
Veröffentlicht: (2026)
von: Rombaut, Benjamin
Veröffentlicht: (2026)
Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG
von: Yu, Boxi, et al.
Veröffentlicht: (2026)
von: Yu, Boxi, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChain
von: Min, Marcus J., et al.
Veröffentlicht: (2023)
von: Min, Marcus J., et al.
Veröffentlicht: (2023)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
A Pattern Language for Resilient Visual Agents
von: Gidey, Habtom Kahsay, et al.
Veröffentlicht: (2026)
von: Gidey, Habtom Kahsay, et al.
Veröffentlicht: (2026)
WasmWalker: Path-based Code Representations for Improved WebAssembly Program Analysis
von: Shirzad, Mohammad Robati, et al.
Veröffentlicht: (2024)
von: Shirzad, Mohammad Robati, et al.
Veröffentlicht: (2024)
Synergy of Large Language Model and Model Driven Engineering for Automated Development of Centralized Vehicular Systems
von: Petrovic, Nenad, et al.
Veröffentlicht: (2024)
von: Petrovic, Nenad, et al.
Veröffentlicht: (2024)
Towards Single-System Illusion in Software-Defined Vehicles -- Automated, AI-Powered Workflow
von: Lebioda, Krzysztof, et al.
Veröffentlicht: (2024)
von: Lebioda, Krzysztof, et al.
Veröffentlicht: (2024)
You Don't Need Public Tests to Generate Correct Code
von: Silva, Kaushitha, et al.
Veröffentlicht: (2026)
von: Silva, Kaushitha, et al.
Veröffentlicht: (2026)
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
von: Yu, Boxi, et al.
Veröffentlicht: (2026)
von: Yu, Boxi, et al.
Veröffentlicht: (2026)
Code Generation for Machine Learning using Model-Driven Engineering and SysML
von: Raedler, Simon, et al.
Veröffentlicht: (2023)
von: Raedler, Simon, et al.
Veröffentlicht: (2023)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
von: Iscan, Mehmet
Veröffentlicht: (2026)
von: Iscan, Mehmet
Veröffentlicht: (2026)
Cognitive Atrophy and Systemic Collapse in AI-Dependent Software Engineering
von: Ginac, Frank
Veröffentlicht: (2026)
von: Ginac, Frank
Veröffentlicht: (2026)
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
von: Gautam, Dhruv, et al.
Veröffentlicht: (2025)
von: Gautam, Dhruv, et al.
Veröffentlicht: (2025)
Large Language Models (LLMs) for Requirements Engineering (RE): A Systematic Literature Review
von: Zadenoori, Mohammad Amin, et al.
Veröffentlicht: (2025)
von: Zadenoori, Mohammad Amin, et al.
Veröffentlicht: (2025)
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
von: Tran, Hung, et al.
Veröffentlicht: (2026)
von: Tran, Hung, et al.
Veröffentlicht: (2026)
EyeLayer: Integrating Human Attention Patterns into LLM-Based Code Summarization
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
AgentPulse: A Continuous Multi-Signal Framework for Evaluating AI Agents in Deployment
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
GenAIOps for GenAI Model-Agility
von: Ueno, Ken, et al.
Veröffentlicht: (2024)
von: Ueno, Ken, et al.
Veröffentlicht: (2024)
ASE-26: a curriculum for agentic software engineering as a discipline
von: Gorsky, Mikael
Veröffentlicht: (2026)
von: Gorsky, Mikael
Veröffentlicht: (2026)
CoverUp: Effective High Coverage Test Generation for Python
von: Pizzorno, Juan Altmayer, et al.
Veröffentlicht: (2024)
von: Pizzorno, Juan Altmayer, et al.
Veröffentlicht: (2024)
LLMDFA: Analyzing Dataflow in Code with Large Language Models
von: Wang, Chengpeng, et al.
Veröffentlicht: (2024)
von: Wang, Chengpeng, et al.
Veröffentlicht: (2024)
AdvFusion: Adapter-based Knowledge Transfer for Code Summarization on Code Language Models
von: Saberi, Iman, et al.
Veröffentlicht: (2023)
von: Saberi, Iman, et al.
Veröffentlicht: (2023)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
Anka: A Domain-Specific Language for Reliable LLM Code Generation
von: Mazrouei, Saif Khalfan Saif Al
Veröffentlicht: (2025)
von: Mazrouei, Saif Khalfan Saif Al
Veröffentlicht: (2025)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
von: Weng, Haojun, et al.
Veröffentlicht: (2026)
von: Weng, Haojun, et al.
Veröffentlicht: (2026)
Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents
von: Tuan, Thanh Luong, et al.
Veröffentlicht: (2026)
von: Tuan, Thanh Luong, et al.
Veröffentlicht: (2026)
State of the Practice for Medical Imaging Software
von: Smith, W. Spencer, et al.
Veröffentlicht: (2024)
von: Smith, W. Spencer, et al.
Veröffentlicht: (2024)
An Exploratory Study on Fine-Tuning Large Language Models for Secure Code Generation
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
A domain-specific language for describing machine learning datasets
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2022)
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2022)
CodeTracer: Towards Traceable Agent States
von: Li, Han, et al.
Veröffentlicht: (2026)
von: Li, Han, et al.
Veröffentlicht: (2026)
Utilization of Pre-trained Language Model for Adapter-based Knowledge Transfer in Software Engineering
von: Saberi, Iman, et al.
Veröffentlicht: (2023)
von: Saberi, Iman, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey
von: Li, Caihua, et al.
Veröffentlicht: (2026) -
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
von: Li, Kenan, et al.
Veröffentlicht: (2026) -
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
von: Dinu, Ion George, et al.
Veröffentlicht: (2026) -
Automated Deep Learning Optimization via DSL-Based Source Code Transformation
von: Wang, Ruixin, et al.
Veröffentlicht: (2024) -
CIFE: Code Instruction-Following Evaluation
von: Gunnu, Sravani, et al.
Veröffentlicht: (2025)