AI Observability for Large Language Model Systems: A Multi-Layer Analysis of Monitoring Approaches from Confidence Calibration to Infrastructure Tracing
Fuente:
arXiv
Saved in:
| Main Author: | Sisodia, Twinkll |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Natural Language to PromQL: A Catalog-Driven Framework with Dynamic Temporal Resolution for Cloud-Native Observability
by: Sisodia, Twinkll
Published: (2026)
by: Sisodia, Twinkll
Published: (2026)
AI Observability for Developer Productivity Tools: Bridging Cost Awareness and Code Quality
by: Bhati, Happy, et al.
Published: (2026)
by: Bhati, Happy, et al.
Published: (2026)
Monitoring and Observability of Machine Learning Systems: Current Practices and Gaps
by: Leest, Joran, et al.
Published: (2025)
by: Leest, Joran, et al.
Published: (2025)
AgentTrace: A Structured Logging Framework for Agent System Observability
by: AlSayyad, Adam, et al.
Published: (2026)
by: AlSayyad, Adam, et al.
Published: (2026)
Fine-grained Approaches for Confidence Calibration of LLMs in Automated Code Revision
by: Lin, Hong Yi, et al.
Published: (2026)
by: Lin, Hong Yi, et al.
Published: (2026)
Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis
by: Rasheed, Zeeshan, et al.
Published: (2024)
by: Rasheed, Zeeshan, et al.
Published: (2024)
Statistical Confidence in Functional Correctness: An Approach for AI Product Functional Correctness Evaluation
by: Albertini, Wallace, et al.
Published: (2026)
by: Albertini, Wallace, et al.
Published: (2026)
Generative AI in Systems Engineering: A Framework for Risk Assessment of Large Language Models
by: Otten, Stefan, et al.
Published: (2026)
by: Otten, Stefan, et al.
Published: (2026)
Large Language Model Evaluation Via Multi AI Agents: Preliminary results
by: Rasheed, Zeeshan, et al.
Published: (2024)
by: Rasheed, Zeeshan, et al.
Published: (2024)
CARGO: A Framework for Confidence-Aware Routing of Large Language Models
by: Barrak, Amine, et al.
Published: (2025)
by: Barrak, Amine, et al.
Published: (2025)
Enhancing COBOL Code Explanations: A Multi-Agents Approach Using Large Language Models
by: Lei, Fangjian, et al.
Published: (2025)
by: Lei, Fangjian, et al.
Published: (2025)
A Survey of using Large Language Models for Generating Infrastructure as Code
by: Srivatsa, Kalahasti Ganesh, et al.
Published: (2024)
by: Srivatsa, Kalahasti Ganesh, et al.
Published: (2024)
Calibration of Large Language Models on Code Summarization
by: Virk, Yuvraj, et al.
Published: (2024)
by: Virk, Yuvraj, et al.
Published: (2024)
TraceLLM: Leveraging Large Language Models with Prompt Engineering for Enhanced Requirements Traceability
by: Alturayeif, Nouf, et al.
Published: (2026)
by: Alturayeif, Nouf, et al.
Published: (2026)
Tracing the Lifecycle of Architecture Technical Debt in Software Systems: A Dependency Approach
by: Sutoyo, Edi, et al.
Published: (2025)
by: Sutoyo, Edi, et al.
Published: (2025)
Sound Concurrent Traces for Online Monitoring Technical Report
by: Soueidi, Chukri, et al.
Published: (2024)
by: Soueidi, Chukri, et al.
Published: (2024)
Victor Calibration (VC): Multi-Pass Confidence Calibration and CP4.3 Governance Stress Test under Round-Table Orchestration
by: Stasiuc, Victor
Published: (2025)
by: Stasiuc, Victor
Published: (2025)
A Large Language Model Approach to Identify Flakiness in C++ Projects
by: Sun, Xin, et al.
Published: (2024)
by: Sun, Xin, et al.
Published: (2024)
A Multi-Language Object-Oriented Programming Benchmark for Large Language Models
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Tracing and Metrics Design Patterns for Monitoring Cloud-native Applications
by: Albuquerque, Carlos, et al.
Published: (2025)
by: Albuquerque, Carlos, et al.
Published: (2025)
LeGEND: A Top-Down Approach to Scenario Generation of Autonomous Driving Systems Assisted by Large Language Models
by: Tang, Shuncheng, et al.
Published: (2024)
by: Tang, Shuncheng, et al.
Published: (2024)
Test Case Generation from Bug Reports via Large Language Models: A Cognitive Layered Evaluation Framework
by: Qureshi, Irtaza Sajid, et al.
Published: (2025)
by: Qureshi, Irtaza Sajid, et al.
Published: (2025)
Does In-IDE Calibration of Large Language Models work at Scale?
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
A Trace-based Approach for Code Safety Analysis
by: Xu, Hui
Published: (2025)
by: Xu, Hui
Published: (2025)
Benchmarking Large Language Models for Multi-Language Software Vulnerability Detection
by: Zhang, Ting, et al.
Published: (2025)
by: Zhang, Ting, et al.
Published: (2025)
AgentSight: System-Level Observability for AI Agents Using eBPF
by: Zheng, Yusheng, et al.
Published: (2025)
by: Zheng, Yusheng, et al.
Published: (2025)
Codified Context: Infrastructure for AI Agents in a Complex Codebase
by: Vasilopoulos, Aristidis
Published: (2026)
by: Vasilopoulos, Aristidis
Published: (2026)
Can Small GenAI Language Models Rival Large Language Models in Understanding Application Behavior?
by: Meymani, Mohammad, et al.
Published: (2025)
by: Meymani, Mohammad, et al.
Published: (2025)
A Learning Method for Symbolic Systems Using Large Language Models
by: Fang, Jian, et al.
Published: (2026)
by: Fang, Jian, et al.
Published: (2026)
ESG Reporting Lifecycle Management with Large Language Models and AI Agents
by: Hoang, Thong, et al.
Published: (2026)
by: Hoang, Thong, et al.
Published: (2026)
Meta-Fair: AI-Assisted Fairness Testing of Large Language Models
by: Romero-Arjona, Miguel, et al.
Published: (2025)
by: Romero-Arjona, Miguel, et al.
Published: (2025)
Towards Synthetic Trace Generation of Modeling Operations using In-Context Learning Approach
by: Muttillo, Vittoriano, et al.
Published: (2024)
by: Muttillo, Vittoriano, et al.
Published: (2024)
A Structured Approach to Safety Case Construction for AI Systems
by: Lee, Sung Une, et al.
Published: (2026)
by: Lee, Sung Une, et al.
Published: (2026)
Semantic-Enhanced Indirect Call Analysis with Large Language Models
by: Cheng, Baijun, et al.
Published: (2024)
by: Cheng, Baijun, et al.
Published: (2024)
Sema Code: Decoupling AI Coding Agents into Programmable, Embeddable Infrastructure
by: Wang, Huacan, et al.
Published: (2026)
by: Wang, Huacan, et al.
Published: (2026)
BinMetric: A Comprehensive Binary Analysis Benchmark for Large Language Models
by: Shang, Xiuwei, et al.
Published: (2025)
by: Shang, Xiuwei, et al.
Published: (2025)
A Tool for In-depth Analysis of Code Execution Reasoning of Large Language Models
by: Liu, Changshu, et al.
Published: (2025)
by: Liu, Changshu, et al.
Published: (2025)
Code Vulnerability Detection: A Comparative Analysis of Emerging Large Language Models
by: Sultana, Shaznin, et al.
Published: (2024)
by: Sultana, Shaznin, et al.
Published: (2024)
Code Digital Twin: A Knowledge Infrastructure for AI-Assisted Complex Software Development
by: Peng, Xin, et al.
Published: (2025)
by: Peng, Xin, et al.
Published: (2025)
Exploring the Potential of Large Language Models in Self-adaptive Systems
by: Li, Jialong, et al.
Published: (2024)
by: Li, Jialong, et al.
Published: (2024)
Similar Items
-
From Natural Language to PromQL: A Catalog-Driven Framework with Dynamic Temporal Resolution for Cloud-Native Observability
by: Sisodia, Twinkll
Published: (2026) -
AI Observability for Developer Productivity Tools: Bridging Cost Awareness and Code Quality
by: Bhati, Happy, et al.
Published: (2026) -
Monitoring and Observability of Machine Learning Systems: Current Practices and Gaps
by: Leest, Joran, et al.
Published: (2025) -
AgentTrace: A Structured Logging Framework for Agent System Observability
by: AlSayyad, Adam, et al.
Published: (2026) -
Fine-grained Approaches for Confidence Calibration of LLMs in Automated Code Revision
by: Lin, Hong Yi, et al.
Published: (2026)