Automated structural testing of LLM-based agents: methods, framework, and case studies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kohl, Jens, Kruse, Otto, Mostafa, Youssef, Luckow, Andre, Schroer, Karsten, Riedl, Thomas, French, Ryan, Katz, David, Luitz, Manuel P., Takher, Tanrajbir, Friedl, Ken E., Laurent-Winter, Céline |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
von: Kohl, Jens, et al.
Veröffentlicht: (2024)
von: Kohl, Jens, et al.
Veröffentlicht: (2024)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
von: Wang, Yuchen, et al.
Veröffentlicht: (2026)
von: Wang, Yuchen, et al.
Veröffentlicht: (2026)
AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
von: Costa, Igor
Veröffentlicht: (2026)
von: Costa, Igor
Veröffentlicht: (2026)
AI Agentic workflows and Enterprise APIs: Adapting API architectures for the age of AI agents
von: Tupe, Vaibhav, et al.
Veröffentlicht: (2025)
von: Tupe, Vaibhav, et al.
Veröffentlicht: (2025)
Exploring Robust Multi-Agent Workflows for Environmental Data Management
von: Guan, Boyuan, et al.
Veröffentlicht: (2026)
von: Guan, Boyuan, et al.
Veröffentlicht: (2026)
Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines
von: Ray, Aninda
Veröffentlicht: (2026)
von: Ray, Aninda
Veröffentlicht: (2026)
PRIMA: Operational Patterns for Resilient Multi-Agent Research with Verifiable Identity and Convergent Feedback
von: Annapureddy, Sasank
Veröffentlicht: (2026)
von: Annapureddy, Sasank
Veröffentlicht: (2026)
Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution
von: Rao, Swanand
Veröffentlicht: (2026)
von: Rao, Swanand
Veröffentlicht: (2026)
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
von: Sellami, Khaled, et al.
Veröffentlicht: (2025)
von: Sellami, Khaled, et al.
Veröffentlicht: (2025)
A Framework for Assessing AI Agent Decisions and Outcomes in AutoML Pipelines
von: Du, Gaoyuan, et al.
Veröffentlicht: (2026)
von: Du, Gaoyuan, et al.
Veröffentlicht: (2026)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
von: Rehan, Tzafrir
Veröffentlicht: (2026)
von: Rehan, Tzafrir
Veröffentlicht: (2026)
Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents
von: Tuan, Thanh Luong, et al.
Veröffentlicht: (2026)
von: Tuan, Thanh Luong, et al.
Veröffentlicht: (2026)
Harnessing Embodied Agents: Runtime Governance for Policy-Constrained Execution
von: Qin, Xue, et al.
Veröffentlicht: (2026)
von: Qin, Xue, et al.
Veröffentlicht: (2026)
Decision Evidence Maturity Model for Agentic AI: A Property-Level Method Specification
von: Solozobov, Oleg
Veröffentlicht: (2026)
von: Solozobov, Oleg
Veröffentlicht: (2026)
Who Writes the Docs in SE 3.0? Agent vs. Human Documentation Pull Requests
von: Yamasaki, Kazuma, et al.
Veröffentlicht: (2026)
von: Yamasaki, Kazuma, et al.
Veröffentlicht: (2026)
AI-Assisted Engineering Should Track the Epistemic Status and Temporal Validity of Architectural Decisions
von: Gilda, Sankalp, et al.
Veröffentlicht: (2026)
von: Gilda, Sankalp, et al.
Veröffentlicht: (2026)
Bounded Autonomy for Enterprise AI: Typed Action Contracts and Consumer-Side Execution
von: Sohail, Sarmad, et al.
Veröffentlicht: (2026)
von: Sohail, Sarmad, et al.
Veröffentlicht: (2026)
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
von: Trooskens, Geert, et al.
Veröffentlicht: (2026)
von: Trooskens, Geert, et al.
Veröffentlicht: (2026)
Knowledge Equivalence in Digital Twins of Intelligent Systems
von: Zhang, Nan, et al.
Veröffentlicht: (2022)
von: Zhang, Nan, et al.
Veröffentlicht: (2022)
BACE: LLM-based Code Generation through Bayesian Anchored Co-Evolution of Code and Test Populations
von: Silva, Kaushitha, et al.
Veröffentlicht: (2026)
von: Silva, Kaushitha, et al.
Veröffentlicht: (2026)
Automated Deep Learning Optimization via DSL-Based Source Code Transformation
von: Wang, Ruixin, et al.
Veröffentlicht: (2024)
von: Wang, Ruixin, et al.
Veröffentlicht: (2024)
BPMN to PDDL: Translating Business Workflows for AI Planning
von: Nie, Jasper, et al.
Veröffentlicht: (2025)
von: Nie, Jasper, et al.
Veröffentlicht: (2025)
A Simple Architecture for Enterprise Large Language Model Applications based on Role based security and Clearance Levels using Retrieval-Augmented Generation or Mixture of Experts
von: Özgür, Atilla, et al.
Veröffentlicht: (2024)
von: Özgür, Atilla, et al.
Veröffentlicht: (2024)
Consensus and Synchronization of Multi-agent Systems over Finite Fields -- Graph Topologies
von: Hengster-Movrić, Kristian, et al.
Veröffentlicht: (2026)
von: Hengster-Movrić, Kristian, et al.
Veröffentlicht: (2026)
Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures
von: Rombaut, Benjamin
Veröffentlicht: (2026)
von: Rombaut, Benjamin
Veröffentlicht: (2026)
TML-Bench: Benchmark for Data Science Agents on Tabular ML Tasks
von: Pinchuk, Mykola
Veröffentlicht: (2026)
von: Pinchuk, Mykola
Veröffentlicht: (2026)
CodeTracer: Towards Traceable Agent States
von: Li, Han, et al.
Veröffentlicht: (2026)
von: Li, Han, et al.
Veröffentlicht: (2026)
Quantitative Analysis of Technical Debt and Pattern Violation in Large Language Model Architectures
von: Slater, Tyler
Veröffentlicht: (2025)
von: Slater, Tyler
Veröffentlicht: (2025)
Emergent Coordination in Multi-Agent Language Models
von: Riedl, Christoph
Veröffentlicht: (2025)
von: Riedl, Christoph
Veröffentlicht: (2025)
Retrieval-Conditioned Topology Selection with Provable Budget Conservation for Multi-Agent Code Generation
von: Talluri, Abhijit, et al.
Veröffentlicht: (2026)
von: Talluri, Abhijit, et al.
Veröffentlicht: (2026)
SAGAI-MID: A Generative AI-Driven Middleware for Dynamic Runtime Interoperability
von: Larsen, Oliver Aleksander, et al.
Veröffentlicht: (2026)
von: Larsen, Oliver Aleksander, et al.
Veröffentlicht: (2026)
Monitoring Agentic Systems Before They're Reliable
von: Boston, Marisa Ferrara, et al.
Veröffentlicht: (2026)
von: Boston, Marisa Ferrara, et al.
Veröffentlicht: (2026)
Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework
von: Zietsman, Christo
Veröffentlicht: (2026)
von: Zietsman, Christo
Veröffentlicht: (2026)
Architectural Patterns for Designing Quantum Artificial Intelligence Systems
von: Klymenko, Mykhailo, et al.
Veröffentlicht: (2024)
von: Klymenko, Mykhailo, et al.
Veröffentlicht: (2024)
ECM Contracts: Contract-Aware, Versioned, and Governable Capability Interfaces for Embodied Agents
von: Qin, Xue, et al.
Veröffentlicht: (2026)
von: Qin, Xue, et al.
Veröffentlicht: (2026)
PyGemini: Unified Software Development towards Maritime Autonomy Systems
von: Vasstein, Kjetil, et al.
Veröffentlicht: (2025)
von: Vasstein, Kjetil, et al.
Veröffentlicht: (2025)
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
von: Calboreanu, Elias
Veröffentlicht: (2026)
von: Calboreanu, Elias
Veröffentlicht: (2026)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
von: Dinu, Ion George, et al.
Veröffentlicht: (2026)
von: Dinu, Ion George, et al.
Veröffentlicht: (2026)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
MLOps with Microservices: A Case Study on the Maritime Domain
von: Ferreira, Renato Cordeiro, et al.
Veröffentlicht: (2025)
von: Ferreira, Renato Cordeiro, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
von: Kohl, Jens, et al.
Veröffentlicht: (2024) -
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
von: Wang, Yuchen, et al.
Veröffentlicht: (2026) -
AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
von: Costa, Igor
Veröffentlicht: (2026) -
AI Agentic workflows and Enterprise APIs: Adapting API architectures for the age of AI agents
von: Tupe, Vaibhav, et al.
Veröffentlicht: (2025) -
Exploring Robust Multi-Agent Workflows for Environmental Data Management
von: Guan, Boyuan, et al.
Veröffentlicht: (2026)