Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
Fuente:
arXiv
Salvato in:
| Autore principale: | Rehan, Tzafrir |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution
di: Rao, Swanand
Pubblicazione: (2026)
di: Rao, Swanand
Pubblicazione: (2026)
AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
di: Costa, Igor
Pubblicazione: (2026)
di: Costa, Igor
Pubblicazione: (2026)
Bounded Autonomy for Enterprise AI: Typed Action Contracts and Consumer-Side Execution
di: Sohail, Sarmad, et al.
Pubblicazione: (2026)
di: Sohail, Sarmad, et al.
Pubblicazione: (2026)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
di: Dinu, Ion George, et al.
Pubblicazione: (2026)
di: Dinu, Ion George, et al.
Pubblicazione: (2026)
CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend
di: Meng, Haoming
Pubblicazione: (2026)
di: Meng, Haoming
Pubblicazione: (2026)
CodeTracer: Towards Traceable Agent States
di: Li, Han, et al.
Pubblicazione: (2026)
di: Li, Han, et al.
Pubblicazione: (2026)
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
di: Trooskens, Geert, et al.
Pubblicazione: (2026)
di: Trooskens, Geert, et al.
Pubblicazione: (2026)
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
di: Sellami, Khaled, et al.
Pubblicazione: (2025)
di: Sellami, Khaled, et al.
Pubblicazione: (2025)
Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework
di: Zietsman, Christo
Pubblicazione: (2026)
di: Zietsman, Christo
Pubblicazione: (2026)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
di: Ahmed, Sheikh Nazib, et al.
Pubblicazione: (2026)
di: Ahmed, Sheikh Nazib, et al.
Pubblicazione: (2026)
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
di: Calboreanu, Elias
Pubblicazione: (2026)
di: Calboreanu, Elias
Pubblicazione: (2026)
Proof of Concept as a First-Class Architectural Decision Instrument
di: Antognolli, Bruno Fernando, et al.
Pubblicazione: (2026)
di: Antognolli, Bruno Fernando, et al.
Pubblicazione: (2026)
AI-Assisted Engineering Should Track the Epistemic Status and Temporal Validity of Architectural Decisions
di: Gilda, Sankalp, et al.
Pubblicazione: (2026)
di: Gilda, Sankalp, et al.
Pubblicazione: (2026)
Automated structural testing of LLM-based agents: methods, framework, and case studies
di: Kohl, Jens, et al.
Pubblicazione: (2026)
di: Kohl, Jens, et al.
Pubblicazione: (2026)
Monitoring Agentic Systems Before They're Reliable
di: Boston, Marisa Ferrara, et al.
Pubblicazione: (2026)
di: Boston, Marisa Ferrara, et al.
Pubblicazione: (2026)
SLEGO: A Collaborative Data Analytics System with LLM Recommender for Diverse Users
di: Ng, Siu Lung, et al.
Pubblicazione: (2024)
di: Ng, Siu Lung, et al.
Pubblicazione: (2024)
AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows
di: Bhardwaj, Varun Pratap
Pubblicazione: (2026)
di: Bhardwaj, Varun Pratap
Pubblicazione: (2026)
AI Agentic workflows and Enterprise APIs: Adapting API architectures for the age of AI agents
di: Tupe, Vaibhav, et al.
Pubblicazione: (2025)
di: Tupe, Vaibhav, et al.
Pubblicazione: (2025)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
di: Wang, Yuchen, et al.
Pubblicazione: (2026)
di: Wang, Yuchen, et al.
Pubblicazione: (2026)
Quantitative Analysis of Technical Debt and Pattern Violation in Large Language Model Architectures
di: Slater, Tyler
Pubblicazione: (2025)
di: Slater, Tyler
Pubblicazione: (2025)
Extending Structural Causal Models for Autonomous Vehicles to Simplify Temporal System Construction & Enable Dynamic Interactions Between Agents
di: Howard, Rhys, et al.
Pubblicazione: (2024)
di: Howard, Rhys, et al.
Pubblicazione: (2024)
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
di: Wu, Jiaqing, et al.
Pubblicazione: (2026)
di: Wu, Jiaqing, et al.
Pubblicazione: (2026)
XARP Tools: An Extended Reality Platform for Humans and AI Agents
di: Caetano, Arthur, et al.
Pubblicazione: (2025)
di: Caetano, Arthur, et al.
Pubblicazione: (2025)
A Framework for Assessing AI Agent Decisions and Outcomes in AutoML Pipelines
di: Du, Gaoyuan, et al.
Pubblicazione: (2026)
di: Du, Gaoyuan, et al.
Pubblicazione: (2026)
Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures
di: Rombaut, Benjamin
Pubblicazione: (2026)
di: Rombaut, Benjamin
Pubblicazione: (2026)
Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents
di: Tuan, Thanh Luong, et al.
Pubblicazione: (2026)
di: Tuan, Thanh Luong, et al.
Pubblicazione: (2026)
The Specification as Quality Gate: Three Hypotheses on AI-Assisted Code Review
di: Zietsman, Christo
Pubblicazione: (2026)
di: Zietsman, Christo
Pubblicazione: (2026)
IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development
di: Moreira, Jasmine
Pubblicazione: (2026)
di: Moreira, Jasmine
Pubblicazione: (2026)
TDAD: Test-Driven Agentic Development - Reducing Code Regressions in AI Coding Agents via Graph-Based Impact Analysis
di: Alonso, Pepe, et al.
Pubblicazione: (2026)
di: Alonso, Pepe, et al.
Pubblicazione: (2026)
The Evolutionary Ecology of Software: Constraints, Innovation, and the AI Disruption
di: Valverde, Sergi, et al.
Pubblicazione: (2025)
di: Valverde, Sergi, et al.
Pubblicazione: (2025)
xML-workFlow: an end-to-end explainable scikit-learn workflow for rapid biomedical experimentation
di: Tran, Khoa A., et al.
Pubblicazione: (2025)
di: Tran, Khoa A., et al.
Pubblicazione: (2025)
Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents
di: Khatchadourian, Raffi
Pubblicazione: (2026)
di: Khatchadourian, Raffi
Pubblicazione: (2026)
Who Writes the Docs in SE 3.0? Agent vs. Human Documentation Pull Requests
di: Yamasaki, Kazuma, et al.
Pubblicazione: (2026)
di: Yamasaki, Kazuma, et al.
Pubblicazione: (2026)
Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery
di: Agarwal, Abhinav
Pubblicazione: (2026)
di: Agarwal, Abhinav
Pubblicazione: (2026)
Review Beats Planning: Dual-Model Interaction Patterns for Code Synthesis
di: Miller, Jan
Pubblicazione: (2026)
di: Miller, Jan
Pubblicazione: (2026)
Towards Observation Lakehouses: Living, Interactive Archives of Software Behavior
di: Kessel, Marcus
Pubblicazione: (2025)
di: Kessel, Marcus
Pubblicazione: (2025)
ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree
di: Koc, Vincent, et al.
Pubblicazione: (2026)
di: Koc, Vincent, et al.
Pubblicazione: (2026)
Generative AI and the Transformation of Software Development Practices
di: Acharya, Vivek
Pubblicazione: (2025)
di: Acharya, Vivek
Pubblicazione: (2025)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
di: Liu, Shunyu, et al.
Pubblicazione: (2025)
di: Liu, Shunyu, et al.
Pubblicazione: (2025)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
di: Bradbury, Jeremy S., et al.
Pubblicazione: (2024)
di: Bradbury, Jeremy S., et al.
Pubblicazione: (2024)
Documenti analoghi
-
Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution
di: Rao, Swanand
Pubblicazione: (2026) -
AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
di: Costa, Igor
Pubblicazione: (2026) -
Bounded Autonomy for Enterprise AI: Typed Action Contracts and Consumer-Side Execution
di: Sohail, Sarmad, et al.
Pubblicazione: (2026) -
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
di: Dinu, Ion George, et al.
Pubblicazione: (2026) -
CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend
di: Meng, Haoming
Pubblicazione: (2026)