ABTest: Behavior-Driven Testing for AI Coding Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Dai, Wuyang, Openja, Moses, Pham, Hung Viet, Uddin, Gias, Yang, Jinqiu, Wang, Song |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Engineering Pitfalls in AI Coding Tools: An Empirical Study of Bugs in Claude Code, Codex, and Gemini CLI
par: Zhang, Ruixin, et autres
Publié: (2026)
par: Zhang, Ruixin, et autres
Publié: (2026)
LLM Assisted Coding with Metamorphic Specification Mutation Agent
par: Akhond, Mostafijur Rahman, et autres
Publié: (2025)
par: Akhond, Mostafijur Rahman, et autres
Publié: (2025)
PAGENT: Learning to Patch Software Engineering Agents
par: Xue, Haoran, et autres
Publié: (2025)
par: Xue, Haoran, et autres
Publié: (2025)
StaAgent: An Agentic Framework for Testing Static Analyzers
par: Nnorom, Elijah, et autres
Publié: (2025)
par: Nnorom, Elijah, et autres
Publié: (2025)
Optimized Log Parsing with Syntactic Modifications
par: Enan, Nafid, et autres
Publié: (2025)
par: Enan, Nafid, et autres
Publié: (2025)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
par: Aleithan, Reem, et autres
Publié: (2024)
par: Aleithan, Reem, et autres
Publié: (2024)
Specification-Driven Code Translation Powered by Large Language Models: How Far Are We?
par: Saha, Soumit Kanti, et autres
Publié: (2024)
par: Saha, Soumit Kanti, et autres
Publié: (2024)
Assessing the Influence of Toxic and Gender Discriminatory Communication on Perceptible Diversity in OSS Projects
par: Sultana, Sayma, et autres
Publié: (2024)
par: Sultana, Sayma, et autres
Publié: (2024)
Evaluating the Environmental Impact of using SLMs and Prompt Engineering for Code Generation
par: Mamun, Md Afif Al, et autres
Publié: (2026)
par: Mamun, Md Afif Al, et autres
Publié: (2026)
An Empirical Study on Bug Severity Estimation using Source Code Metrics and Static Analysis
par: Mashhadi, Ehsan, et autres
Publié: (2022)
par: Mashhadi, Ehsan, et autres
Publié: (2022)
FairFLRep: Fairness aware fault localization and repair of Deep Neural Networks
par: Openja, Moses, et autres
Publié: (2025)
par: Openja, Moses, et autres
Publié: (2025)
Perception-Guided Fuzzing for Simulated Scenario-Based Testing of Autonomous Driving Systems
par: Pham, Tri Minh Triet, et autres
Publié: (2024)
par: Pham, Tri Minh Triet, et autres
Publié: (2024)
Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
par: Salimian, Sina, et autres
Publié: (2025)
par: Salimian, Sina, et autres
Publié: (2025)
SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection
par: Ahmed, Md Basim Uddin, et autres
Publié: (2025)
par: Ahmed, Md Basim Uddin, et autres
Publié: (2025)
ChatGPT Inaccuracy Mitigation during Technical Report Understanding: Are We There Yet?
par: Tamanna, Salma Begum, et autres
Publié: (2024)
par: Tamanna, Salma Begum, et autres
Publié: (2024)
Checker Bug Detection and Repair in Deep Learning Libraries
par: Harzevili, Nima Shiri, et autres
Publié: (2024)
par: Harzevili, Nima Shiri, et autres
Publié: (2024)
Bias Unveiled: Investigating Social Bias in LLM-Generated Code
par: Ling, Lin, et autres
Publié: (2024)
par: Ling, Lin, et autres
Publié: (2024)
Selection of Prompt Engineering Techniques for Code Generation through Predicting Code Complexity
par: Wang, Chung-Yu, et autres
Publié: (2024)
par: Wang, Chung-Yu, et autres
Publié: (2024)
Reputation Gaming in Stack Overflow
par: Mazloomzadeh, Iren, et autres
Publié: (2021)
par: Mazloomzadeh, Iren, et autres
Publié: (2021)
BLAgent: Agentic RAG for File-Level Bug Localization
par: Mamun, Md Afif Al, et autres
Publié: (2026)
par: Mamun, Md Afif Al, et autres
Publié: (2026)
CFCEval: Evaluating Security Aspects in Code Generated by Large Language Models
par: Cheng, Cheng, et autres
Publié: (2025)
par: Cheng, Cheng, et autres
Publié: (2025)
Tracing Stereotypes in Pre-trained Transformers: From Biased Neurons to Fairer Models
par: Voria, Gianmario, et autres
Publié: (2026)
par: Voria, Gianmario, et autres
Publié: (2026)
Tracking the Evolution of Static Code Warnings: the State-of-the-Art and a Better Approach
par: Li, Junjie, et autres
Publié: (2022)
par: Li, Junjie, et autres
Publié: (2022)
ADPerf: Investigating and Testing Performance in Autonomous Driving Systems
par: Pham, Tri Minh-Triet, et autres
Publié: (2025)
par: Pham, Tri Minh-Triet, et autres
Publié: (2025)
LLM For Loop Invariant Generation and Fixing: How Far Are We?
par: Akhond, Mostafijur Rahman, et autres
Publié: (2025)
par: Akhond, Mostafijur Rahman, et autres
Publié: (2025)
BabelCoder: Agentic Code Translation with Specification Alignment
par: Rabbi, Fazle, et autres
Publié: (2025)
par: Rabbi, Fazle, et autres
Publié: (2025)
Secure-Instruct: An Automated Pipeline for Synthesizing Instruction-Tuning Datasets Using LLMs for Secure Code Generation
par: Li, Junjie, et autres
Publié: (2025)
par: Li, Junjie, et autres
Publié: (2025)
A Systematic Mapping Study of Crowd Knowledge Enhanced Software Engineering Research Using Stack Overflow
par: Tanzil, Minaoar, et autres
Publié: (2024)
par: Tanzil, Minaoar, et autres
Publié: (2024)
A Multi-Language Perspective on the Robustness of LLM Code Generation
par: Rabbi, Fazle, et autres
Publié: (2025)
par: Rabbi, Fazle, et autres
Publié: (2025)
On the Robustness Evaluation of 3D Obstacle Detection Against Specifications in Autonomous Driving
par: Pham, Tri Minh Triet, et autres
Publié: (2024)
par: Pham, Tri Minh Triet, et autres
Publié: (2024)
"How do people decide?": A Model for Software Library Selection
par: Tanzil, Minaoar Hossain, et autres
Publié: (2024)
par: Tanzil, Minaoar Hossain, et autres
Publié: (2024)
Secret Leak Detection in Software Issue Reports using LLMs: A Comprehensive Evaluation
par: Ahmed, Sadif, et autres
Publié: (2024)
par: Ahmed, Sadif, et autres
Publié: (2024)
A Large-Scale Empirical Study of COVID-19 Contact Tracing Mobile App Reviews
par: Parisa, Sifat Ishmam, et autres
Publié: (2024)
par: Parisa, Sifat Ishmam, et autres
Publié: (2024)
Consistency Meets Verification: Enhancing Test Generation Quality in Large Language Models Without Ground-Truth Solutions
par: Taherkhani, Hamed, et autres
Publié: (2026)
par: Taherkhani, Hamed, et autres
Publié: (2026)
Beyond Translation Accuracy: Addressing False Failures in LLM-Based Code Translation
par: Rabbi, Fazle, et autres
Publié: (2026)
par: Rabbi, Fazle, et autres
Publié: (2026)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
par: Xu, Yisen, et autres
Publié: (2026)
par: Xu, Yisen, et autres
Publié: (2026)
Automated Prompt Engineering for Cost-Effective Code Generation Using Evolutionary Algorithm
par: Taherkhani, Hamed, et autres
Publié: (2024)
par: Taherkhani, Hamed, et autres
Publié: (2024)
Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
par: Abdollahi, Mohammad, et autres
Publié: (2025)
par: Abdollahi, Mohammad, et autres
Publié: (2025)
RGFL: Reasoning Guided Fault Localization for Automated Program Repair Using Large Language Models
par: Sepidband, Melika, et autres
Publié: (2026)
par: Sepidband, Melika, et autres
Publié: (2026)
An empirical study of testing machine learning in the wild
par: Openja, Moses, et autres
Publié: (2023)
par: Openja, Moses, et autres
Publié: (2023)
Documents similaires
-
Engineering Pitfalls in AI Coding Tools: An Empirical Study of Bugs in Claude Code, Codex, and Gemini CLI
par: Zhang, Ruixin, et autres
Publié: (2026) -
LLM Assisted Coding with Metamorphic Specification Mutation Agent
par: Akhond, Mostafijur Rahman, et autres
Publié: (2025) -
PAGENT: Learning to Patch Software Engineering Agents
par: Xue, Haoran, et autres
Publié: (2025) -
StaAgent: An Agentic Framework for Testing Static Analyzers
par: Nnorom, Elijah, et autres
Publié: (2025) -
Optimized Log Parsing with Syntactic Modifications
par: Enan, Nafid, et autres
Publié: (2025)