Evaluating LLM Agents on Automated Software Analysis Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Bouzenia, Islem, Cadar, Cristian, Pradel, Michael |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
por: Bouzenia, Islem, et al.
Publicado: (2025)
por: Bouzenia, Islem, et al.
Publicado: (2025)
You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects
por: Bouzenia, Islem, et al.
Publicado: (2024)
por: Bouzenia, Islem, et al.
Publicado: (2024)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
por: Bouzenia, Islem, et al.
Publicado: (2024)
por: Bouzenia, Islem, et al.
Publicado: (2024)
CodeCureAgent: Automatic Classification and Repair of Static Analysis Warnings
por: Joos, Pascal, et al.
Publicado: (2025)
por: Joos, Pascal, et al.
Publicado: (2025)
DyPyBench: A Benchmark of Executable Python Software
por: Bouzenia, Islem, et al.
Publicado: (2024)
por: Bouzenia, Islem, et al.
Publicado: (2024)
Issue2Test: Generating Reproducing Test Cases from Issue Reports
por: Nashid, Noor, et al.
Publicado: (2025)
por: Nashid, Noor, et al.
Publicado: (2025)
PatchGuru: Patch Oracle Inference from Natural Language Artifacts with Large Language Models
por: Le-Cong, Thanh, et al.
Publicado: (2026)
por: Le-Cong, Thanh, et al.
Publicado: (2026)
AgentStepper: Interactive Debugging of Software Development Agents
por: Hutter, Robert, et al.
Publicado: (2026)
por: Hutter, Robert, et al.
Publicado: (2026)
Understanding API Usage and Testing: An Empirical Study of C Libraries
por: Zaki, Ahmed, et al.
Publicado: (2025)
por: Zaki, Ahmed, et al.
Publicado: (2025)
De-Hallucinator: Mitigating LLM Hallucinations in Code Generation Tasks via Iterative Grounding
por: Eghbali, Aryaz, et al.
Publicado: (2024)
por: Eghbali, Aryaz, et al.
Publicado: (2024)
Software Security Analysis in 2030 and Beyond: A Research Roadmap
por: Böhme, Marcel, et al.
Publicado: (2024)
por: Böhme, Marcel, et al.
Publicado: (2024)
Artisan: Agentic Artifact Evaluation
por: Baek, Doehyun, et al.
Publicado: (2026)
por: Baek, Doehyun, et al.
Publicado: (2026)
Testora: Using Natural Language Intent to Detect Behavioral Regressions
por: Pradel, Michael
Publicado: (2025)
por: Pradel, Michael
Publicado: (2025)
CodeMapper: A Language-Agnostic Approach to Mapping Code Regions Across Commits
por: Hu, Huimin, et al.
Publicado: (2025)
por: Hu, Huimin, et al.
Publicado: (2025)
Agentic AI Software Engineers: Programming with Trust
por: Roychoudhury, Abhik, et al.
Publicado: (2025)
por: Roychoudhury, Abhik, et al.
Publicado: (2025)
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
por: Badertdinov, Ibragim, et al.
Publicado: (2025)
por: Badertdinov, Ibragim, et al.
Publicado: (2025)
From Bugs to Benchmarks: A Comprehensive Survey of Software Defect Datasets
por: Zhu, Hao-Nan, et al.
Publicado: (2025)
por: Zhu, Hao-Nan, et al.
Publicado: (2025)
Humans Integrate, Agents Fix: How Agent-Authored Pull Requests Are Referenced in Practice
por: Khemissi, Islem, et al.
Publicado: (2026)
por: Khemissi, Islem, et al.
Publicado: (2026)
Analyzing Quantum Programs with LintQ: A Static Analysis Framework for Qiskit
por: Paltenghi, Matteo, et al.
Publicado: (2023)
por: Paltenghi, Matteo, et al.
Publicado: (2023)
Automated Personnel Selection for Software Engineers Using LLM-Based Profile Evaluation
por: Karim, Ahmed Akib Jawad, et al.
Publicado: (2024)
por: Karim, Ahmed Akib Jawad, et al.
Publicado: (2024)
CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building
por: Yu, Zhengmin, et al.
Publicado: (2025)
por: Yu, Zhengmin, et al.
Publicado: (2025)
A Survey on Testing and Analysis of Quantum Software
por: Paltenghi, Matteo, et al.
Publicado: (2024)
por: Paltenghi, Matteo, et al.
Publicado: (2024)
A Comprehensive Empirical Evaluation of Agent Frameworks on Code-centric Software Engineering Tasks
por: Yin, Zhuowen, et al.
Publicado: (2025)
por: Yin, Zhuowen, et al.
Publicado: (2025)
RippleGUItester: Change-Aware Exploratory Testing
por: Su, Yanqi, et al.
Publicado: (2026)
por: Su, Yanqi, et al.
Publicado: (2026)
Names Are All You Need: Effective and Safe Regression Test Selection for Python
por: Wang, You, et al.
Publicado: (2026)
por: Wang, You, et al.
Publicado: (2026)
Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
por: Wang, You, et al.
Publicado: (2025)
por: Wang, You, et al.
Publicado: (2025)
ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution
por: Gröninger, Lars, et al.
Publicado: (2024)
por: Gröninger, Lars, et al.
Publicado: (2024)
LLM-Agents Driven Automated Simulation Testing and Analysis of small Uncrewed Aerial Systems
por: Duvvuru, Venkata Sai Aswath, et al.
Publicado: (2025)
por: Duvvuru, Venkata Sai Aswath, et al.
Publicado: (2025)
NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition
por: Deng, Le, et al.
Publicado: (2025)
por: Deng, Le, et al.
Publicado: (2025)
BASFuzz: Towards Robustness Evaluation of LLM-based NLP Software via Automated Fuzz Testing
por: Xiao, Mingxuan, et al.
Publicado: (2025)
por: Xiao, Mingxuan, et al.
Publicado: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
por: Fandina, Ora Nova, et al.
Publicado: (2025)
por: Fandina, Ora Nova, et al.
Publicado: (2025)
Treefix: Enabling Execution with a Tree of Prefixes
por: Souza, Beatriz, et al.
Publicado: (2025)
por: Souza, Beatriz, et al.
Publicado: (2025)
Beyond pip install: Evaluating LLM Agents for the Automated Installation of Python Projects
por: Milliken, Louis, et al.
Publicado: (2024)
por: Milliken, Louis, et al.
Publicado: (2024)
Agent-Based Software Artifact Evaluation
por: Wu, Zhaonan, et al.
Publicado: (2026)
por: Wu, Zhaonan, et al.
Publicado: (2026)
LLM Agents for Automated Dependency Upgrades
por: Tawosi, Vali, et al.
Publicado: (2025)
por: Tawosi, Vali, et al.
Publicado: (2025)
Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents
por: Xiang, Jiahong, et al.
Publicado: (2026)
por: Xiang, Jiahong, et al.
Publicado: (2026)
LLM-Based Repair of Static Nullability Errors
por: Karimipour, Nima, et al.
Publicado: (2025)
por: Karimipour, Nima, et al.
Publicado: (2025)
Assessing the Robustness of LLM-based NLP Software via Automated Testing
por: Xiao, Mingxuan, et al.
Publicado: (2024)
por: Xiao, Mingxuan, et al.
Publicado: (2024)
Knowledge-Based Multi-Agent Framework for Automated Software Architecture Design
por: Zhang, Yiran, et al.
Publicado: (2025)
por: Zhang, Yiran, et al.
Publicado: (2025)
FlyCatcher: Neural Inference of Runtime Checkers from Tests
por: Souza, Beatriz, et al.
Publicado: (2026)
por: Souza, Beatriz, et al.
Publicado: (2026)
Ejemplares similares
-
Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
por: Bouzenia, Islem, et al.
Publicado: (2025) -
You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects
por: Bouzenia, Islem, et al.
Publicado: (2024) -
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
por: Bouzenia, Islem, et al.
Publicado: (2024) -
CodeCureAgent: Automatic Classification and Repair of Static Analysis Warnings
por: Joos, Pascal, et al.
Publicado: (2025) -
DyPyBench: A Benchmark of Executable Python Software
por: Bouzenia, Islem, et al.
Publicado: (2024)