CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend
Fuente:
arXiv
Saved in:
| Main Author: | Meng, Haoming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Monitoring Agentic Systems Before They're Reliable
by: Boston, Marisa Ferrara, et al.
Published: (2026)
by: Boston, Marisa Ferrara, et al.
Published: (2026)
Automated structural testing of LLM-based agents: methods, framework, and case studies
by: Kohl, Jens, et al.
Published: (2026)
by: Kohl, Jens, et al.
Published: (2026)
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
by: Calboreanu, Elias
Published: (2026)
by: Calboreanu, Elias
Published: (2026)
Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution
by: Rao, Swanand
Published: (2026)
by: Rao, Swanand
Published: (2026)
CodeTracer: Towards Traceable Agent States
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
by: Rehan, Tzafrir
Published: (2026)
by: Rehan, Tzafrir
Published: (2026)
ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree
by: Koc, Vincent, et al.
Published: (2026)
by: Koc, Vincent, et al.
Published: (2026)
Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents
by: Tuan, Thanh Luong, et al.
Published: (2026)
by: Tuan, Thanh Luong, et al.
Published: (2026)
Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery
by: Agarwal, Abhinav
Published: (2026)
by: Agarwal, Abhinav
Published: (2026)
Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response
by: Drammeh, Philip
Published: (2025)
by: Drammeh, Philip
Published: (2025)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
by: Dinu, Ion George, et al.
Published: (2026)
by: Dinu, Ion George, et al.
Published: (2026)
Proof of Concept as a First-Class Architectural Decision Instrument
by: Antognolli, Bruno Fernando, et al.
Published: (2026)
by: Antognolli, Bruno Fernando, et al.
Published: (2026)
AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
by: Costa, Igor
Published: (2026)
by: Costa, Igor
Published: (2026)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
by: Wang, Yuchen, et al.
Published: (2026)
by: Wang, Yuchen, et al.
Published: (2026)
xML-workFlow: an end-to-end explainable scikit-learn workflow for rapid biomedical experimentation
by: Tran, Khoa A., et al.
Published: (2025)
by: Tran, Khoa A., et al.
Published: (2025)
The Evolutionary Ecology of Software: Constraints, Innovation, and the AI Disruption
by: Valverde, Sergi, et al.
Published: (2025)
by: Valverde, Sergi, et al.
Published: (2025)
Bounded Autonomy for Enterprise AI: Typed Action Contracts and Consumer-Side Execution
by: Sohail, Sarmad, et al.
Published: (2026)
by: Sohail, Sarmad, et al.
Published: (2026)
AI Agentic workflows and Enterprise APIs: Adapting API architectures for the age of AI agents
by: Tupe, Vaibhav, et al.
Published: (2025)
by: Tupe, Vaibhav, et al.
Published: (2025)
ClawVM: Harness-Managed Virtual Memory for Stateful Tool-Using LLM Agents
by: Rafique, Mofasshara, et al.
Published: (2026)
by: Rafique, Mofasshara, et al.
Published: (2026)
Toolsuite for Implementing Multiagent Systems Based on Communication Protocols
by: Chopra, Amit K., et al.
Published: (2025)
by: Chopra, Amit K., et al.
Published: (2025)
CodeCRDT: Observation-Driven Coordination for Multi-Agent LLM Code Generation
by: Pugachev, Sergey
Published: (2025)
by: Pugachev, Sergey
Published: (2025)
Technical Debt Management: The Road Ahead for Successful Software Delivery
by: Avgeriou, Paris, et al.
Published: (2024)
by: Avgeriou, Paris, et al.
Published: (2024)
A Systematic Review of Digital Twin-Driven Predictive Maintenance in Industrial Engineering: Taxonomy, Architectural Elements, and Future Research Directions
by: Ismail, Leila, et al.
Published: (2025)
by: Ismail, Leila, et al.
Published: (2025)
Making a Pipeline Production-Ready: Challenges and Lessons Learned in the Healthcare Domain
by: Lawand, Daniel Angelo Esteves, et al.
Published: (2025)
by: Lawand, Daniel Angelo Esteves, et al.
Published: (2025)
SPIRA: Building an Intelligent System for Respiratory Insufficiency Detection
by: Ferreira, Renato Cordeiro, et al.
Published: (2025)
by: Ferreira, Renato Cordeiro, et al.
Published: (2025)
Automated Deep Learning Optimization via DSL-Based Source Code Transformation
by: Wang, Ruixin, et al.
Published: (2024)
by: Wang, Ruixin, et al.
Published: (2024)
Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework
by: Zietsman, Christo
Published: (2026)
by: Zietsman, Christo
Published: (2026)
Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents
by: Khatchadourian, Raffi
Published: (2026)
by: Khatchadourian, Raffi
Published: (2026)
Who Writes the Docs in SE 3.0? Agent vs. Human Documentation Pull Requests
by: Yamasaki, Kazuma, et al.
Published: (2026)
by: Yamasaki, Kazuma, et al.
Published: (2026)
Good modelling software practices
by: Lemmen, Carsten, et al.
Published: (2024)
by: Lemmen, Carsten, et al.
Published: (2024)
A Framework for Assessing AI Agent Decisions and Outcomes in AutoML Pipelines
by: Du, Gaoyuan, et al.
Published: (2026)
by: Du, Gaoyuan, et al.
Published: (2026)
BONSAI: A Mixed-Initiative Workspace for Human-AI Co-Development of Visual Analytics Applications
by: Spinner, Thilo, et al.
Published: (2026)
by: Spinner, Thilo, et al.
Published: (2026)
When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape
by: Mitchell, Richard Joseph
Published: (2026)
by: Mitchell, Richard Joseph
Published: (2026)
A Pattern Language for Resilient Visual Agents
by: Gidey, Habtom Kahsay, et al.
Published: (2026)
by: Gidey, Habtom Kahsay, et al.
Published: (2026)
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
by: Sellami, Khaled, et al.
Published: (2025)
by: Sellami, Khaled, et al.
Published: (2025)
SLEGO: A Collaborative Data Analytics System with LLM Recommender for Diverse Users
by: Ng, Siu Lung, et al.
Published: (2024)
by: Ng, Siu Lung, et al.
Published: (2024)
Agentic Business Process Management: Practitioner Perspectives on Agent Governance in Business Processes
by: Vu, Hoang, et al.
Published: (2025)
by: Vu, Hoang, et al.
Published: (2025)
What You See Is What It Does: A Structural Pattern for Legible Software
by: Meng, Eagon, et al.
Published: (2025)
by: Meng, Eagon, et al.
Published: (2025)
Evaluation of MQTT Bridge Architectures in a Cross-Organizational Context
by: Lima, Keila, et al.
Published: (2025)
by: Lima, Keila, et al.
Published: (2025)
Separating Intelligence from Execution: A Workflow Engine for the Model Context Protocol
by: Parmar, Abhinav Singh
Published: (2026)
by: Parmar, Abhinav Singh
Published: (2026)
Similar Items
-
Monitoring Agentic Systems Before They're Reliable
by: Boston, Marisa Ferrara, et al.
Published: (2026) -
Automated structural testing of LLM-based agents: methods, framework, and case studies
by: Kohl, Jens, et al.
Published: (2026) -
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
by: Calboreanu, Elias
Published: (2026) -
Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution
by: Rao, Swanand
Published: (2026) -
CodeTracer: Towards Traceable Agent States
by: Li, Han, et al.
Published: (2026)