Saved in:
| Main Authors: | Dougherty, Quinn, von Hippel, Max, Shackleton, Hazel, Dodds, Mike |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2606.01008 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agentic Property-Based Testing: Finding Bugs Across the Python Ecosystem
by: Maaz, Muhammad, et al.
Published: (2025)
by: Maaz, Muhammad, et al.
Published: (2025)
Proving the Coding Interview: A Benchmark for Formally Verified Code Generation
by: Dougherty, Quinn, et al.
Published: (2025)
by: Dougherty, Quinn, et al.
Published: (2025)
Transforming Software Development: Evaluating the Efficiency and Challenges of GitHub Copilot in Real-World Projects
by: Pandey, Ruchika, et al.
Published: (2024)
by: Pandey, Ruchika, et al.
Published: (2024)
PBT-Bench: Benchmarking AI Agents on Property-Based Testing
by: Jing, Lucas, et al.
Published: (2026)
by: Jing, Lucas, et al.
Published: (2026)
Challenges in Testing Large Language Model Based Software: A Faceted Taxonomy
by: Dobslaw, Felix, et al.
Published: (2025)
by: Dobslaw, Felix, et al.
Published: (2025)
Agentic Harness for Real-World Compilers
by: Zheng, Yingwei, et al.
Published: (2026)
by: Zheng, Yingwei, et al.
Published: (2026)
SCoGen: Scenario-Centric Graph-Based Synthesis of Real-World Code Problems
by: Yao, Xifeng, et al.
Published: (2025)
by: Yao, Xifeng, et al.
Published: (2025)
Harden and Catch for Just-in-Time Assured LLM-Based Software Testing: Open Research Challenges
by: Harman, Mark, et al.
Published: (2025)
by: Harman, Mark, et al.
Published: (2025)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
by: Mündler, Niels, et al.
Published: (2024)
by: Mündler, Niels, et al.
Published: (2024)
Deep Learning Library Testing: Definition, Methods and Challenges
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
SWE-Universe: Scale Real-World Verifiable Environments to Millions
by: Chen, Mouxiang, et al.
Published: (2026)
by: Chen, Mouxiang, et al.
Published: (2026)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
by: Liang, Jiarong, et al.
Published: (2026)
by: Liang, Jiarong, et al.
Published: (2026)
An Empirical Study of Proactive Coding Assistants in Real-World Software Development
by: Li, Lehui, et al.
Published: (2026)
by: Li, Lehui, et al.
Published: (2026)
Evaluating the Effectiveness of LLMs in Fixing Maintainability Issues in Real-World Projects
by: Nunes, Henrique, et al.
Published: (2025)
by: Nunes, Henrique, et al.
Published: (2025)
Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
by: Ma, Wei, et al.
Published: (2025)
by: Ma, Wei, et al.
Published: (2025)
Reasoning-Based Software Testing
by: Giamattei, Luca, et al.
Published: (2023)
by: Giamattei, Luca, et al.
Published: (2023)
Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis
by: Akli, Amal, et al.
Published: (2026)
by: Akli, Amal, et al.
Published: (2026)
CIRCLE: A Framework for Evaluating AI from a Real-World Lens
by: Schwartz, Reva, et al.
Published: (2026)
by: Schwartz, Reva, et al.
Published: (2026)
CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
by: Roy, Monoshi Kumar, et al.
Published: (2025)
by: Roy, Monoshi Kumar, et al.
Published: (2025)
RepoMasterEval: Evaluating Code Completion via Real-World Repositories
by: Wu, Qinyun, et al.
Published: (2024)
by: Wu, Qinyun, et al.
Published: (2024)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
by: Gao, Zeyu, et al.
Published: (2025)
by: Gao, Zeyu, et al.
Published: (2025)
LoCoML: A Framework for Real-World ML Inference Pipelines
by: Maddireddy, Kritin, et al.
Published: (2025)
by: Maddireddy, Kritin, et al.
Published: (2025)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
by: Ma, Jeffrey Jian, et al.
Published: (2025)
by: Ma, Jeffrey Jian, et al.
Published: (2025)
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
by: Li, Chenxin, et al.
Published: (2026)
by: Li, Chenxin, et al.
Published: (2026)
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
Bridging the Gap between Real-world and Synthetic Images for Testing Autonomous Driving Systems
by: Amini, Mohammad Hossein, et al.
Published: (2024)
by: Amini, Mohammad Hossein, et al.
Published: (2024)
System Test Case Design from Requirements Specifications: Insights and Challenges of Using ChatGPT
by: Bhatia, Shreya, et al.
Published: (2024)
by: Bhatia, Shreya, et al.
Published: (2024)
Inferring Code Correctness from Specification
by: Florian, Tambon, et al.
Published: (2026)
by: Florian, Tambon, et al.
Published: (2026)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
by: Han, Tingxu, et al.
Published: (2026)
by: Han, Tingxu, et al.
Published: (2026)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
by: Xu, Qiao, et al.
Published: (2026)
by: Xu, Qiao, et al.
Published: (2026)
PerfBench: Can Agents Resolve Real-World Performance Bugs?
by: Garg, Spandan, et al.
Published: (2025)
by: Garg, Spandan, et al.
Published: (2025)
LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities
by: Tang, Yongjian, et al.
Published: (2026)
by: Tang, Yongjian, et al.
Published: (2026)
LLM-Based Robustness Testing of Microservice Applications: An Empirical Study
by: Tigulla, Hrushitha Goud, et al.
Published: (2026)
by: Tigulla, Hrushitha Goud, et al.
Published: (2026)
LLM-Based Automated Diagnosis Of Integration Test Failures At Google
by: Ziftci, Celal, et al.
Published: (2026)
by: Ziftci, Celal, et al.
Published: (2026)
Evaluating LLM-Based Test Generation Under Software Evolution
by: Haroon, Sabaat, et al.
Published: (2026)
by: Haroon, Sabaat, et al.
Published: (2026)
SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Resolving Real-World Bugs
by: Pham, Minh V. T., et al.
Published: (2025)
by: Pham, Minh V. T., et al.
Published: (2025)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarios
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
Deploying Geospatial Foundation Models in the Real World: Lessons from WorldCereal
by: Butsko, Christina, et al.
Published: (2025)
by: Butsko, Christina, et al.
Published: (2025)
Similar Items
-
Agentic Property-Based Testing: Finding Bugs Across the Python Ecosystem
by: Maaz, Muhammad, et al.
Published: (2025) -
Proving the Coding Interview: A Benchmark for Formally Verified Code Generation
by: Dougherty, Quinn, et al.
Published: (2025) -
Transforming Software Development: Evaluating the Efficiency and Challenges of GitHub Copilot in Real-World Projects
by: Pandey, Ruchika, et al.
Published: (2024) -
PBT-Bench: Benchmarking AI Agents on Property-Based Testing
by: Jing, Lucas, et al.
Published: (2026) -
Challenges in Testing Large Language Model Based Software: A Faceted Taxonomy
by: Dobslaw, Felix, et al.
Published: (2025)