AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows
Fuente:
arXiv
Guardado en:
| Autor principal: | Bhardwaj, Varun Pratap |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
por: Rehan, Tzafrir
Publicado: (2026)
por: Rehan, Tzafrir
Publicado: (2026)
Reasoning Provenance for Autonomous AI Agents: Structured Behavioral Analytics Beyond State Checkpoints and Execution Traces
por: Vispute, Neelmani, et al.
Publicado: (2026)
por: Vispute, Neelmani, et al.
Publicado: (2026)
Fuzzing the brain: Automated stress testing for the safety of ML-driven neurostimulation
por: Downing, Mara, et al.
Publicado: (2025)
por: Downing, Mara, et al.
Publicado: (2025)
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
por: Li, Meiziniu, et al.
Publicado: (2024)
por: Li, Meiziniu, et al.
Publicado: (2024)
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
por: Li, Meiziniu, et al.
Publicado: (2022)
por: Li, Meiziniu, et al.
Publicado: (2022)
Social, Legal, Ethical, Empathetic and Cultural Norm Operationalisation for AI Agents
por: Calinescu, Radu, et al.
Publicado: (2026)
por: Calinescu, Radu, et al.
Publicado: (2026)
Context-Augmented Code Generation: How Product Context Improves AI Coding Agent Decision Compliance by 49%
por: Dillon, Drew, et al.
Publicado: (2026)
por: Dillon, Drew, et al.
Publicado: (2026)
PICKLES: a Natural Language Framework for Requirement Specification and Model-Based Testing
por: Rodríguez, María Belén, et al.
Publicado: (2026)
por: Rodríguez, María Belén, et al.
Publicado: (2026)
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
por: Ahmed, Syed Yusuf, et al.
Publicado: (2026)
por: Ahmed, Syed Yusuf, et al.
Publicado: (2026)
Demystifying the Silence of Correctness Bugs in PyTorch Compiler
por: Li, Meiziniu, et al.
Publicado: (2026)
por: Li, Meiziniu, et al.
Publicado: (2026)
CodeTracer: Towards Traceable Agent States
por: Li, Han, et al.
Publicado: (2026)
por: Li, Han, et al.
Publicado: (2026)
TDAD: Test-Driven Agentic Development - Reducing Code Regressions in AI Coding Agents via Graph-Based Impact Analysis
por: Alonso, Pepe, et al.
Publicado: (2026)
por: Alonso, Pepe, et al.
Publicado: (2026)
LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB
por: Bekmyradov, Vekil, et al.
Publicado: (2026)
por: Bekmyradov, Vekil, et al.
Publicado: (2026)
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
por: Terragni, Valerio
Publicado: (2026)
por: Terragni, Valerio
Publicado: (2026)
Towards Explainable Test Case Prioritisation with Learning-to-Rank Models
por: Ramírez, Aurora, et al.
Publicado: (2024)
por: Ramírez, Aurora, et al.
Publicado: (2024)
Multi-Agent Code Verification via Information Theory
por: Rajan, Shreshth
Publicado: (2025)
por: Rajan, Shreshth
Publicado: (2025)
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
por: More, Riddhi, et al.
Publicado: (2025)
por: More, Riddhi, et al.
Publicado: (2025)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
por: Hu, Yuelin, et al.
Publicado: (2026)
por: Hu, Yuelin, et al.
Publicado: (2026)
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
por: More, Riddhi, et al.
Publicado: (2025)
por: More, Riddhi, et al.
Publicado: (2025)
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
por: Chang, Hung-Fu, et al.
Publicado: (2025)
por: Chang, Hung-Fu, et al.
Publicado: (2025)
Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
Synthesizing Test Cases for Narrowing Specification Candidates
por: Cunha, Alcino, et al.
Publicado: (2025)
por: Cunha, Alcino, et al.
Publicado: (2025)
Validating Formal Specifications with LLM-generated Test Cases
por: Cunha, Alcino, et al.
Publicado: (2025)
por: Cunha, Alcino, et al.
Publicado: (2025)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
por: Bradbury, Jeremy S., et al.
Publicado: (2024)
por: Bradbury, Jeremy S., et al.
Publicado: (2024)
Abstain and Validate: A Dual-LLM Policy for Reducing Noise in Agentic Program Repair
por: Cambronero, José, et al.
Publicado: (2025)
por: Cambronero, José, et al.
Publicado: (2025)
Orion: Fuzzing Workflow Automation
por: Bazalii, Max, et al.
Publicado: (2025)
por: Bazalii, Max, et al.
Publicado: (2025)
AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code
por: Parris, William M.
Publicado: (2026)
por: Parris, William M.
Publicado: (2026)
Behavioral Fingerprints for LLM Endpoint Stability and Identity
por: Leshin, Jonah, et al.
Publicado: (2026)
por: Leshin, Jonah, et al.
Publicado: (2026)
The Future of AI-Driven Software Engineering
por: Terragni, Valerio, et al.
Publicado: (2024)
por: Terragni, Valerio, et al.
Publicado: (2024)
PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows
por: Kawamura, Kazuki, et al.
Publicado: (2026)
por: Kawamura, Kazuki, et al.
Publicado: (2026)
Bounded Autonomy for Enterprise AI: Typed Action Contracts and Consumer-Side Execution
por: Sohail, Sarmad, et al.
Publicado: (2026)
por: Sohail, Sarmad, et al.
Publicado: (2026)
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
por: Wu, Jiaqing, et al.
Publicado: (2026)
por: Wu, Jiaqing, et al.
Publicado: (2026)
On the Soundness and Consistency of LLM Agents for Executing Test Cases Written in Natural Language
por: Salva, Sébastien, et al.
Publicado: (2025)
por: Salva, Sébastien, et al.
Publicado: (2025)
On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning Implementations
por: Hundal, Rajdeep Singh, et al.
Publicado: (2025)
por: Hundal, Rajdeep Singh, et al.
Publicado: (2025)
Who is Introducing the Failure? Automatically Attributing Failures of Multi-Agent Systems via Spectrum Analysis
por: Ge, Yu, et al.
Publicado: (2025)
por: Ge, Yu, et al.
Publicado: (2025)
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
por: Calboreanu, Elias
Publicado: (2026)
por: Calboreanu, Elias
Publicado: (2026)
Experience with GitHub Copilot for Developer Productivity at Zoominfo
por: Bakal, Gal, et al.
Publicado: (2025)
por: Bakal, Gal, et al.
Publicado: (2025)
Talk is Cheap, Logic is Hard: Benchmarking LLMs on Post-Condition Formalization
por: Prasetya, I. S. W. B., et al.
Publicado: (2026)
por: Prasetya, I. S. W. B., et al.
Publicado: (2026)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
por: Ahmed, Sheikh Nazib, et al.
Publicado: (2026)
por: Ahmed, Sheikh Nazib, et al.
Publicado: (2026)
InterEvo-TR: Interactive Evolutionary Test Generation With Readability Assessment
por: Delgado-Pérez, Pedro, et al.
Publicado: (2024)
por: Delgado-Pérez, Pedro, et al.
Publicado: (2024)
Ejemplares similares
-
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
por: Rehan, Tzafrir
Publicado: (2026) -
Reasoning Provenance for Autonomous AI Agents: Structured Behavioral Analytics Beyond State Checkpoints and Execution Traces
por: Vispute, Neelmani, et al.
Publicado: (2026) -
Fuzzing the brain: Automated stress testing for the safety of ML-driven neurostimulation
por: Downing, Mara, et al.
Publicado: (2025) -
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
por: Li, Meiziniu, et al.
Publicado: (2024) -
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
por: Li, Meiziniu, et al.
Publicado: (2022)