Monitoring Agentic Systems Before They're Reliable
Fuente:
arXiv
Guardado en:
| Autores principales: | Boston, Marisa Ferrara, Hanson, Glen, Georgala, Effi, Hudgens, JD, Frase, Heather |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution
por: Rao, Swanand
Publicado: (2026)
por: Rao, Swanand
Publicado: (2026)
Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents
por: Tuan, Thanh Luong, et al.
Publicado: (2026)
por: Tuan, Thanh Luong, et al.
Publicado: (2026)
Automated structural testing of LLM-based agents: methods, framework, and case studies
por: Kohl, Jens, et al.
Publicado: (2026)
por: Kohl, Jens, et al.
Publicado: (2026)
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
por: Calboreanu, Elias
Publicado: (2026)
por: Calboreanu, Elias
Publicado: (2026)
Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework
por: Zietsman, Christo
Publicado: (2026)
por: Zietsman, Christo
Publicado: (2026)
CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend
por: Meng, Haoming
Publicado: (2026)
por: Meng, Haoming
Publicado: (2026)
Decision Evidence Maturity Model for Agentic AI: A Property-Level Method Specification
por: Solozobov, Oleg
Publicado: (2026)
por: Solozobov, Oleg
Publicado: (2026)
CodeTracer: Towards Traceable Agent States
por: Li, Han, et al.
Publicado: (2026)
por: Li, Han, et al.
Publicado: (2026)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
por: Rehan, Tzafrir
Publicado: (2026)
por: Rehan, Tzafrir
Publicado: (2026)
AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
por: Costa, Igor
Publicado: (2026)
por: Costa, Igor
Publicado: (2026)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
por: Wang, Yuchen, et al.
Publicado: (2026)
por: Wang, Yuchen, et al.
Publicado: (2026)
AI Agentic workflows and Enterprise APIs: Adapting API architectures for the age of AI agents
por: Tupe, Vaibhav, et al.
Publicado: (2025)
por: Tupe, Vaibhav, et al.
Publicado: (2025)
Proof of Concept as a First-Class Architectural Decision Instrument
por: Antognolli, Bruno Fernando, et al.
Publicado: (2026)
por: Antognolli, Bruno Fernando, et al.
Publicado: (2026)
Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response
por: Drammeh, Philip
Publicado: (2025)
por: Drammeh, Philip
Publicado: (2025)
Architectural Transformations and Emerging Verification Demands in AI-Enabled Cyber-Physical Systems
por: Yusuf, Hadiza Umar, et al.
Publicado: (2025)
por: Yusuf, Hadiza Umar, et al.
Publicado: (2025)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
por: Dinu, Ion George, et al.
Publicado: (2026)
por: Dinu, Ion George, et al.
Publicado: (2026)
BACE: LLM-based Code Generation through Bayesian Anchored Co-Evolution of Code and Test Populations
por: Silva, Kaushitha, et al.
Publicado: (2026)
por: Silva, Kaushitha, et al.
Publicado: (2026)
Toward Causal-Visual Programming: Enhancing Agentic Reasoning in Low-Code Environments
por: Xu, Jiexi, et al.
Publicado: (2025)
por: Xu, Jiexi, et al.
Publicado: (2025)
xML-workFlow: an end-to-end explainable scikit-learn workflow for rapid biomedical experimentation
por: Tran, Khoa A., et al.
Publicado: (2025)
por: Tran, Khoa A., et al.
Publicado: (2025)
MFH: A Multi-faceted Heuristic Algorithm Selection Approach for Software Verification
por: Su, Jie, et al.
Publicado: (2025)
por: Su, Jie, et al.
Publicado: (2025)
Automated Deep Learning Optimization via DSL-Based Source Code Transformation
por: Wang, Ruixin, et al.
Publicado: (2024)
por: Wang, Ruixin, et al.
Publicado: (2024)
Who Writes the Docs in SE 3.0? Agent vs. Human Documentation Pull Requests
por: Yamasaki, Kazuma, et al.
Publicado: (2026)
por: Yamasaki, Kazuma, et al.
Publicado: (2026)
Good modelling software practices
por: Lemmen, Carsten, et al.
Publicado: (2024)
por: Lemmen, Carsten, et al.
Publicado: (2024)
Agentic Business Process Management: Practitioner Perspectives on Agent Governance in Business Processes
por: Vu, Hoang, et al.
Publicado: (2025)
por: Vu, Hoang, et al.
Publicado: (2025)
Technical Debt Management: The Road Ahead for Successful Software Delivery
por: Avgeriou, Paris, et al.
Publicado: (2024)
por: Avgeriou, Paris, et al.
Publicado: (2024)
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
por: Sellami, Khaled, et al.
Publicado: (2025)
por: Sellami, Khaled, et al.
Publicado: (2025)
ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree
por: Koc, Vincent, et al.
Publicado: (2026)
por: Koc, Vincent, et al.
Publicado: (2026)
From Everything-is-a-File to Files-Are-All-You-Need: How Unix Philosophy Informs the Design of Agentic AI Systems
por: Piskala, Deepak Babu
Publicado: (2026)
por: Piskala, Deepak Babu
Publicado: (2026)
Bounded Autonomy for Enterprise AI: Typed Action Contracts and Consumer-Side Execution
por: Sohail, Sarmad, et al.
Publicado: (2026)
por: Sohail, Sarmad, et al.
Publicado: (2026)
The Evolutionary Ecology of Software: Constraints, Innovation, and the AI Disruption
por: Valverde, Sergi, et al.
Publicado: (2025)
por: Valverde, Sergi, et al.
Publicado: (2025)
ABACUS: A FinOps Service for Cloud Cost Optimization
por: Deochake, Saurabh
Publicado: (2024)
por: Deochake, Saurabh
Publicado: (2024)
A Framework for Assessing AI Agent Decisions and Outcomes in AutoML Pipelines
por: Du, Gaoyuan, et al.
Publicado: (2026)
por: Du, Gaoyuan, et al.
Publicado: (2026)
Exploring Robust Multi-Agent Workflows for Environmental Data Management
por: Guan, Boyuan, et al.
Publicado: (2026)
por: Guan, Boyuan, et al.
Publicado: (2026)
Quantitative Analysis of Technical Debt and Pattern Violation in Large Language Model Architectures
por: Slater, Tyler
Publicado: (2025)
por: Slater, Tyler
Publicado: (2025)
Nidus: Externalized Reasoning for AI-Assisted Engineering
por: Gorinevski, Danil
Publicado: (2026)
por: Gorinevski, Danil
Publicado: (2026)
Adaptable TeaStore: A Choreographic Approach
por: De Palma, Giuseppe, et al.
Publicado: (2025)
por: De Palma, Giuseppe, et al.
Publicado: (2025)
CodeCRDT: Observation-Driven Coordination for Multi-Agent LLM Code Generation
por: Pugachev, Sergey
Publicado: (2025)
por: Pugachev, Sergey
Publicado: (2025)
Separating Intelligence from Execution: A Workflow Engine for the Model Context Protocol
por: Parmar, Abhinav Singh
Publicado: (2026)
por: Parmar, Abhinav Singh
Publicado: (2026)
When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape
por: Mitchell, Richard Joseph
Publicado: (2026)
por: Mitchell, Richard Joseph
Publicado: (2026)
An LLM-assisted approach to designing software architectures using ADD
por: Cervantes, Humberto, et al.
Publicado: (2025)
por: Cervantes, Humberto, et al.
Publicado: (2025)
Ejemplares similares
-
Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution
por: Rao, Swanand
Publicado: (2026) -
Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents
por: Tuan, Thanh Luong, et al.
Publicado: (2026) -
Automated structural testing of LLM-based agents: methods, framework, and case studies
por: Kohl, Jens, et al.
Publicado: (2026) -
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
por: Calboreanu, Elias
Publicado: (2026) -
Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework
por: Zietsman, Christo
Publicado: (2026)