Evaluating LLM-Based Test Generation Under Software Evolution
Fuente:
arXiv
Guardado en:
| Autores principales: | Haroon, Sabaat, Khan, Mohammad Taha, Gulzar, Muhammad Ali |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models
por: Haroon, Sabaat, et al.
Publicado: (2025)
por: Haroon, Sabaat, et al.
Publicado: (2025)
The Future of Software Testing: AI-Powered Test Case Generation and Validation
por: Baqar, Mohammad, et al.
Publicado: (2024)
por: Baqar, Mohammad, et al.
Publicado: (2024)
The Rise of Agentic Testing: Multi-Agent Systems for Robust Software Quality Assurance
por: Naqvi, Saba, et al.
Publicado: (2026)
por: Naqvi, Saba, et al.
Publicado: (2026)
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
por: Chen, Zhi, et al.
Publicado: (2026)
por: Chen, Zhi, et al.
Publicado: (2026)
Can LLM Generate Regression Tests for Software Commits?
por: Liu, Jing, et al.
Publicado: (2025)
por: Liu, Jing, et al.
Publicado: (2025)
Agentic AI in 6G Software Businesses: A Layered Maturity Model
por: Zohaib, Muhammad, et al.
Publicado: (2025)
por: Zohaib, Muhammad, et al.
Publicado: (2025)
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
por: Hu, Ruida, et al.
Publicado: (2026)
por: Hu, Ruida, et al.
Publicado: (2026)
OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
por: Imran, Mia Mohammad, et al.
Publicado: (2025)
por: Imran, Mia Mohammad, et al.
Publicado: (2025)
Reasoning-Based Software Testing
por: Giamattei, Luca, et al.
Publicado: (2023)
por: Giamattei, Luca, et al.
Publicado: (2023)
ProToken: Token-Level Attribution for Federated Large Language Models
por: Gill, Waris, et al.
Publicado: (2026)
por: Gill, Waris, et al.
Publicado: (2026)
Efficient Test Data Generation for MC/DC with OCL and Search
por: Sartaj, Hassan, et al.
Publicado: (2024)
por: Sartaj, Hassan, et al.
Publicado: (2024)
Breaking Barriers in Software Testing: The Power of AI-Driven Automation
por: Naqvi, Saba, et al.
Publicado: (2025)
por: Naqvi, Saba, et al.
Publicado: (2025)
LLM Test Generation via Iterative Hybrid Program Analysis
por: Gu, Sijia, et al.
Publicado: (2025)
por: Gu, Sijia, et al.
Publicado: (2025)
Harden and Catch for Just-in-Time Assured LLM-Based Software Testing: Open Research Challenges
por: Harman, Mark, et al.
Publicado: (2025)
por: Harman, Mark, et al.
Publicado: (2025)
Towards Specification-Driven LLM-Based Generation of Embedded Automotive Software
por: Patil, Minal Suresh, et al.
Publicado: (2024)
por: Patil, Minal Suresh, et al.
Publicado: (2024)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
por: Agarwal, Anisha, et al.
Publicado: (2024)
por: Agarwal, Anisha, et al.
Publicado: (2024)
EvoClaw: Evaluating AI Agents on Continuous Software Evolution
por: Deng, Gangda, et al.
Publicado: (2026)
por: Deng, Gangda, et al.
Publicado: (2026)
EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
por: Kaliyev, Alibek T., et al.
Publicado: (2026)
por: Kaliyev, Alibek T., et al.
Publicado: (2026)
On Simulation-Guided LLM-based Code Generation for Safe Autonomous Driving Software
por: Nouri, Ali, et al.
Publicado: (2025)
por: Nouri, Ali, et al.
Publicado: (2025)
LADs: Leveraging LLMs for AI-Driven DevOps
por: Khan, Ahmad Faraz, et al.
Publicado: (2025)
por: Khan, Ahmad Faraz, et al.
Publicado: (2025)
Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling
por: Trae Research Team, et al.
Publicado: (2025)
por: Trae Research Team, et al.
Publicado: (2025)
A Metamorphic Testing Approach to Diagnosing Memorization in LLM-Based Program Repair
por: De Koning, Milan, et al.
Publicado: (2026)
por: De Koning, Milan, et al.
Publicado: (2026)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
por: Fandina, Ora Nova, et al.
Publicado: (2025)
por: Fandina, Ora Nova, et al.
Publicado: (2025)
Data and Context Matter: Towards Generalizing AI-based Software Vulnerability Detection
por: Safdar, Rijha, et al.
Publicado: (2025)
por: Safdar, Rijha, et al.
Publicado: (2025)
Call-Chain-Aware LLM-Based Test Generation for Java Projects
por: Wang, Guancheng, et al.
Publicado: (2026)
por: Wang, Guancheng, et al.
Publicado: (2026)
Static Program Analysis Guided LLM Based Unit Test Generation
por: Roychowdhury, Sujoy, et al.
Publicado: (2025)
por: Roychowdhury, Sujoy, et al.
Publicado: (2025)
The Potential of LLMs in Automating Software Testing: From Generation to Reporting
por: Sherifi, Betim, et al.
Publicado: (2024)
por: Sherifi, Betim, et al.
Publicado: (2024)
AgentForge: Execution-Grounded Multi-Agent LLM Framework for Autonomous Software Engineering
por: Kumar, Rajesh, et al.
Publicado: (2026)
por: Kumar, Rajesh, et al.
Publicado: (2026)
Beyond Isolated Tasks: A Framework for Evaluating Coding Agents on Sequential Software Evolution
por: Shastry, KN Ajay, et al.
Publicado: (2026)
por: Shastry, KN Ajay, et al.
Publicado: (2026)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
por: Pan, Zhiyuan, et al.
Publicado: (2025)
por: Pan, Zhiyuan, et al.
Publicado: (2025)
LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities
por: Tang, Yongjian, et al.
Publicado: (2026)
por: Tang, Yongjian, et al.
Publicado: (2026)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
por: Fan, Zhiyu, et al.
Publicado: (2025)
por: Fan, Zhiyu, et al.
Publicado: (2025)
LLM-Based Test Case Generation in DBMS through Monte Carlo Tree Search
por: Chen, Yujia, et al.
Publicado: (2026)
por: Chen, Yujia, et al.
Publicado: (2026)
Verifying LLM-Generated Code in the Context of Software Verification with Ada/SPARK
por: Cramer, Marcos, et al.
Publicado: (2025)
por: Cramer, Marcos, et al.
Publicado: (2025)
Challenges in Testing Large Language Model Based Software: A Faceted Taxonomy
por: Dobslaw, Felix, et al.
Publicado: (2025)
por: Dobslaw, Felix, et al.
Publicado: (2025)
Generating Privacy Stories From Software Documentation
por: Baldwin, Wilder, et al.
Publicado: (2025)
por: Baldwin, Wilder, et al.
Publicado: (2025)
Fuzzy Inference System for Test Case Prioritization in Software Testing
por: Karatayev, Aron, et al.
Publicado: (2024)
por: Karatayev, Aron, et al.
Publicado: (2024)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
por: Cui, Yi
Publicado: (2025)
por: Cui, Yi
Publicado: (2025)
Automated System-level Testing of Unmanned Aerial Systems
por: Sartaj, Hassan, et al.
Publicado: (2024)
por: Sartaj, Hassan, et al.
Publicado: (2024)
On the Adoption of AI Coding Agents in Open-source Android and iOS Development
por: Khan, Muhammad Ahmad, et al.
Publicado: (2026)
por: Khan, Muhammad Ahmad, et al.
Publicado: (2026)
Ejemplares similares
-
Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models
por: Haroon, Sabaat, et al.
Publicado: (2025) -
The Future of Software Testing: AI-Powered Test Case Generation and Validation
por: Baqar, Mohammad, et al.
Publicado: (2024) -
The Rise of Agentic Testing: Multi-Agent Systems for Robust Software Quality Assurance
por: Naqvi, Saba, et al.
Publicado: (2026) -
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
por: Chen, Zhi, et al.
Publicado: (2026) -
Can LLM Generate Regression Tests for Software Commits?
por: Liu, Jing, et al.
Publicado: (2025)