An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Hasan, Mohammed Mehedi, Li, Hao, Fallahzadeh, Emad, Rajbahadur, Gopi Krishnan, Adams, Bram, Hassan, Ahmed E. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
by: Hasan, Mohammed Mehedi, et al.
Published: (2026)
by: Hasan, Mohammed Mehedi, et al.
Published: (2026)
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
by: Jewitt, James, et al.
Published: (2025)
by: Jewitt, James, et al.
Published: (2025)
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
by: Jewitt, James, et al.
Published: (2026)
by: Jewitt, James, et al.
Published: (2026)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
Data Quality Antipatterns for Software Analytics
by: Bhatia, Aaditya, et al.
Published: (2024)
by: Bhatia, Aaditya, et al.
Published: (2024)
Assessing and Improving the Representativeness of Code Generation Benchmarks Using Knowledge Units (KUs) of Programming Languages -- An Empirical Study
by: Ahasanuzzaman, Md, et al.
Published: (2026)
by: Ahasanuzzaman, Md, et al.
Published: (2026)
A State-of-the-practice Release-readiness Checklist for Generative AI-based Software Products
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks
by: Ebrahimi, Amir M., et al.
Published: (2026)
by: Ebrahimi, Amir M., et al.
Published: (2026)
Agentic Refactoring: An Empirical Study of AI Coding Agents
by: Horikawa, Kosei, et al.
Published: (2025)
by: Horikawa, Kosei, et al.
Published: (2025)
HAFix: History-Augmented Large Language Models for Bug Fixing
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
Reformulating Regression Test Suite Optimization using Quantum Annealing -- an Empirical Study
by: Trovato, Antonio, et al.
Published: (2024)
by: Trovato, Antonio, et al.
Published: (2024)
Building an Open AIBOM Standard in the Wild
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2025)
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2025)
Studying the Impact of TensorFlow and PyTorch Bindings on Machine Learning Software Quality
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Utilizing Composer Packages to Accelerate Laravel-Based Project Development Among Students: A Pedagogical and Practical Framework
by: Wahid, Rohaizah Abdul, et al.
Published: (2025)
by: Wahid, Rohaizah Abdul, et al.
Published: (2025)
AgenticSZZ: Temporal Knowledge Graph-Guided Agentic Bug-Inducing Commit Identification
by: Shi, Yu, et al.
Published: (2026)
by: Shi, Yu, et al.
Published: (2026)
Clawed and Dangerous: Can We Trust Open Agentic Systems?
by: Chen, Shiping, et al.
Published: (2026)
by: Chen, Shiping, et al.
Published: (2026)
Agent READMEs: An Empirical Study of Context Files for Agentic Coding
by: Chatlatanagulchai, Worawalan, et al.
Published: (2025)
by: Chatlatanagulchai, Worawalan, et al.
Published: (2025)
From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2024)
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2024)
Quantum Software Architecture Framework (QSAF): A Component-Based Framework for Designing Hybrid Quantum-Classical Systems
by: Kiwelekar, Arvind W., et al.
Published: (2026)
by: Kiwelekar, Arvind W., et al.
Published: (2026)
A Curated List of Open-source Software-only Energy Efficiency Measurement Tools: A GitHub Mining Study
by: Cannizza, Manuela Bechara, et al.
Published: (2026)
by: Cannizza, Manuela Bechara, et al.
Published: (2026)
OrQstrator: An AI-Powered Framework for Advanced Quantum Circuit Optimization
by: Baird, Laura, et al.
Published: (2025)
by: Baird, Laura, et al.
Published: (2025)
A Semantic Framework for Patient Digital Twins in Chronic Care
by: Elgammal, Amal, et al.
Published: (2025)
by: Elgammal, Amal, et al.
Published: (2025)
LLM4VV: Evaluating Cutting-Edge LLMs for Generation and Evaluation of Directive-Based Parallel Programming Model Compiler Tests
by: Sollenberger, Zachariah, et al.
Published: (2025)
by: Sollenberger, Zachariah, et al.
Published: (2025)
A Context-Driven Approach for Co-Auditing Smart Contracts with The Support of GPT-4 code interpreter
by: Bouafif, Mohamed Salah, et al.
Published: (2024)
by: Bouafif, Mohamed Salah, et al.
Published: (2024)
A Construction-Phase Digital Twin Framework for Quality Assurance and Decision Support in Civil Infrastructure Projects
by: Islam, Md Asiful, et al.
Published: (2026)
by: Islam, Md Asiful, et al.
Published: (2026)
HAFixAgent: History-Aware Program Repair Agent
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
PyEncode: An Open-Source Library for Structured Quantum State Preparation
by: Suresh, Krishnan, et al.
Published: (2026)
by: Suresh, Krishnan, et al.
Published: (2026)
An Empirical Study of Self-Admitted Technical Debt in Machine Learning Software
by: Bhatia, Aaditya, et al.
Published: (2023)
by: Bhatia, Aaditya, et al.
Published: (2023)
Do AI Coding Agents Log Like Humans? An Empirical Study
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2026)
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2026)
SpecPylot: Python Specification Generation using Large Language Models
by: Ayon, Ragib Shahariar, et al.
Published: (2026)
by: Ayon, Ragib Shahariar, et al.
Published: (2026)
Automating Detection and Root-Cause Analysis of Flaky Tests in Quantum Software
by: Sivaloganathan, Janakan, et al.
Published: (2026)
by: Sivaloganathan, Janakan, et al.
Published: (2026)
Implementing AI Bill of Materials (AI BOM) with SPDX 3.0: A Comprehensive Guide to Creating AI and Dataset Bill of Materials
by: Bennet, Karen, et al.
Published: (2025)
by: Bennet, Karen, et al.
Published: (2025)
Human-Certified Module Repositories for the AI Age
by: Enyedi, Szilárd
Published: (2026)
by: Enyedi, Szilárd
Published: (2026)
AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
by: Rana, Manik, et al.
Published: (2025)
by: Rana, Manik, et al.
Published: (2025)
An Empirical Study on Code Review Activity Prediction and Its Impact in Practice
by: Olewicki, Doriane, et al.
Published: (2024)
by: Olewicki, Doriane, et al.
Published: (2024)
AI Simulation by Digital Twins: Systematic Survey, Reference Framework, and Mapping to a Standardized Architecture
by: Liu, Xiaoran, et al.
Published: (2025)
by: Liu, Xiaoran, et al.
Published: (2025)
Holon Programming Model -- A Software-Defined Approach for System of Systems
by: Ashfaq, Muhammad, et al.
Published: (2024)
by: Ashfaq, Muhammad, et al.
Published: (2024)
TwinArch: A Digital Twin Reference Architecture
by: Somma, Alessandra, et al.
Published: (2025)
by: Somma, Alessandra, et al.
Published: (2025)
Process Analytics -- Data-driven Business Process Management
by: Stierle, Matthias, et al.
Published: (2025)
by: Stierle, Matthias, et al.
Published: (2025)
Similar Items
-
Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
by: Hasan, Mohammed Mehedi, et al.
Published: (2025) -
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
by: Hasan, Mohammed Mehedi, et al.
Published: (2026) -
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
by: Jewitt, James, et al.
Published: (2025) -
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
by: Jewitt, James, et al.
Published: (2026) -
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)