Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Rombaut, Benjamin, Masoumzadeh, Sogol, Vasilevski, Kirill, Lin, Dayi, Hassan, Ahmed E. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Detecting Vulnerabilities from Issue Reports for Internet-of-Things
by: Masoumzadeh, Sogol
Published: (2025)
by: Masoumzadeh, Sogol
Published: (2025)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
by: Fan, Zhiyu, et al.
Published: (2025)
by: Fan, Zhiyu, et al.
Published: (2025)
Towards Reliable Generation of Executable Workflows by Foundation Models
by: Masoumzadeh, Sogol, et al.
Published: (2025)
by: Masoumzadeh, Sogol, et al.
Published: (2025)
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
by: Vasilevski, Kirill, et al.
Published: (2025)
by: Vasilevski, Kirill, et al.
Published: (2025)
Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement
by: Gallaba, Keheliya, et al.
Published: (2025)
by: Gallaba, Keheliya, et al.
Published: (2025)
Engineering AI Judge Systems
by: Lin, Jiahuei, et al.
Published: (2024)
by: Lin, Jiahuei, et al.
Published: (2024)
Data Quality Antipatterns for Software Analytics
by: Bhatia, Aaditya, et al.
Published: (2024)
by: Bhatia, Aaditya, et al.
Published: (2024)
Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
Rethinking Software Engineering in the Foundation Model Era: From Task-Driven AI Copilots to Goal-Driven AI Pair Programmers
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2024)
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2024)
Real-time Adapting Routing (RAR): Improving Efficiency Through Continuous Learning in Software Powered by Layered Foundation Models
by: Vasilevski, Kirill, et al.
Published: (2024)
by: Vasilevski, Kirill, et al.
Published: (2024)
AgentTrace: A Structured Logging Framework for Agent System Observability
by: AlSayyad, Adam, et al.
Published: (2026)
by: AlSayyad, Adam, et al.
Published: (2026)
Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures
by: Rombaut, Benjamin
Published: (2026)
by: Rombaut, Benjamin
Published: (2026)
Agentic Software Engineering: Foundational Pillars and a Research Roadmap
by: Hassan, Ahmed E., et al.
Published: (2025)
by: Hassan, Ahmed E., et al.
Published: (2025)
Leveraging the Crowd for Dependency Management: An Empirical Study on the Dependabot Compatibility Score
by: Rombaut, Benjamin, et al.
Published: (2024)
by: Rombaut, Benjamin, et al.
Published: (2024)
HAFixAgent: History-Aware Program Repair Agent
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
AIDev: Studying AI Coding Agents on GitHub
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
AgentForge: Execution-Grounded Multi-Agent LLM Framework for Autonomous Software Engineering
by: Kumar, Rajesh, et al.
Published: (2026)
by: Kumar, Rajesh, et al.
Published: (2026)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
by: Oliva, Gustavo A., et al.
Published: (2025)
by: Oliva, Gustavo A., et al.
Published: (2025)
OmniLLP: Enhancing LLM-based Log Level Prediction with Context-Aware Retrieval
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
Rethinking Software Engineering in the Foundation Model Era: A Curated Catalogue of Challenges in the Development of Trustworthy FMware
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
Software Performance Engineering for Foundation Model-Powered Software
by: Zhang, Haoxiang, et al.
Published: (2024)
by: Zhang, Haoxiang, et al.
Published: (2024)
Keeping Deep Learning Models in Check: A History-Based Approach to Mitigate Overfitting
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
AgentGit: A Version Control Framework for Reliable and Scalable LLM-Powered Multi-Agent Systems
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
A Self-Healing Framework for Reliable LLM-Based Autonomous Agents
by: Jeong, Cheonsu, et al.
Published: (2026)
by: Jeong, Cheonsu, et al.
Published: (2026)
OrcaLoca: An LLM Agent Framework for Software Issue Localization
by: Yu, Zhongming, et al.
Published: (2025)
by: Yu, Zhongming, et al.
Published: (2025)
Vul-R2: A Reasoning LLM for Automated Vulnerability Repair
by: Wen, Xin-Cheng, et al.
Published: (2025)
by: Wen, Xin-Cheng, et al.
Published: (2025)
ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
by: Tawosi, Vali, et al.
Published: (2025)
by: Tawosi, Vali, et al.
Published: (2025)
RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring
by: Oueslati, Khouloud, et al.
Published: (2025)
by: Oueslati, Khouloud, et al.
Published: (2025)
AutoP2C: An LLM-Based Agent Framework for Code Repository Generation from Multimodal Content in Academic Papers
by: Lin, Zijie, et al.
Published: (2025)
by: Lin, Zijie, et al.
Published: (2025)
Z-Space: A Multi-Agent Tool Orchestration Framework for Enterprise-Grade LLM Automation
by: He, Qingsong, et al.
Published: (2025)
by: He, Qingsong, et al.
Published: (2025)
MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution
by: Tao, Wei, et al.
Published: (2024)
by: Tao, Wei, et al.
Published: (2024)
LLM Agents for Interactive Exploration of Historical Cadastre Data: Framework and Application to Venice
by: Karch, Tristan, et al.
Published: (2025)
by: Karch, Tristan, et al.
Published: (2025)
Software Engineering and Foundation Models: Insights from Industry Blogs Using a Jury of Foundation Models
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
by: Son, Ha Min, et al.
Published: (2025)
by: Son, Ha Min, et al.
Published: (2025)
FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning
by: Ding, Haoran, et al.
Published: (2026)
by: Ding, Haoran, et al.
Published: (2026)
MAO: A Framework for Process Model Generation with Multi-Agent Orchestration
by: Lin, Leilei, et al.
Published: (2024)
by: Lin, Leilei, et al.
Published: (2024)
AgenticTCAD: A LLM-based Multi-Agent Framework for Automated TCAD Code Generation and Device Optimization
by: Fan, Guangxi, et al.
Published: (2025)
by: Fan, Guangxi, et al.
Published: (2025)
Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling
by: Trae Research Team, et al.
Published: (2025)
by: Trae Research Team, et al.
Published: (2025)
A State-of-the-practice Release-readiness Checklist for Generative AI-based Software Products
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
Similar Items
-
Detecting Vulnerabilities from Issue Reports for Internet-of-Things
by: Masoumzadeh, Sogol
Published: (2025) -
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
by: Fan, Zhiyu, et al.
Published: (2025) -
Towards Reliable Generation of Executable Workflows by Foundation Models
by: Masoumzadeh, Sogol, et al.
Published: (2025) -
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
by: Vasilevski, Kirill, et al.
Published: (2025) -
Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement
by: Gallaba, Keheliya, et al.
Published: (2025)