Measuring the Unmeasurable: Markov Chain Reliability for LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Tran-Truong, Phat T., Le, Xuan-Bach |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring the Reasoning Depth of Small Language Models in Software Architecture: A Multidimensional Evaluation Framework Towards Software Engineering 2.0
by: Vo, Ha, et al.
Published: (2026)
by: Vo, Ha, et al.
Published: (2026)
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
by: Le-Cong, Thanh, et al.
Published: (2024)
by: Le-Cong, Thanh, et al.
Published: (2024)
Unlocking LLM Repair Capabilities Through Cross-Language Translation and Multi-Agent Refinement
by: Luo, Wenqiang, et al.
Published: (2025)
by: Luo, Wenqiang, et al.
Published: (2025)
Teamwork makes the dream work: LLMs-Based Agents for GitHub README.MD Summarization
by: Nguyen, Duc S. H., et al.
Published: (2025)
by: Nguyen, Duc S. H., et al.
Published: (2025)
VisualCoder: Guiding Large Language Models in Code Execution with Fine-grained Multimodal Chain-of-Thought Reasoning
by: Le, Cuong Chi, et al.
Published: (2024)
by: Le, Cuong Chi, et al.
Published: (2024)
Verify-Gated Completion as Admission Control in a Governed Multi-Agent Runtime: A Bounded Architecture Case Study
by: Nguyen, Hai-Duong, et al.
Published: (2026)
by: Nguyen, Hai-Duong, et al.
Published: (2026)
Adversarial Attacks on Code Models with Discriminative Graph Patterns
by: Nguyen, Thanh-Dat, et al.
Published: (2023)
by: Nguyen, Thanh-Dat, et al.
Published: (2023)
Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond
by: Le-Anh, Minh, et al.
Published: (2026)
by: Le-Anh, Minh, et al.
Published: (2026)
A Self-Healing Framework for Reliable LLM-Based Autonomous Agents
by: Jeong, Cheonsu, et al.
Published: (2026)
by: Jeong, Cheonsu, et al.
Published: (2026)
When Fine-Tuning LLMs Meets Data Privacy: An Empirical Study of Federated Learning in LLM-Based Program Repair
by: Luo, Wenqiang, et al.
Published: (2024)
by: Luo, Wenqiang, et al.
Published: (2024)
Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency
by: Ashrafi, Nazmus, et al.
Published: (2025)
by: Ashrafi, Nazmus, et al.
Published: (2025)
A Survey of LLM-based Automated Program Repair: Taxonomies, Design Paradigms, and Applications
by: Yang, Boyang, et al.
Published: (2025)
by: Yang, Boyang, et al.
Published: (2025)
The Dual-State Architecture for Reliable LLM Agents
by: Thompson, Matthew
Published: (2025)
by: Thompson, Matthew
Published: (2025)
Exploiting LLM Agent Supply Chains via Payload-less Skills
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
PatchGuru: Patch Oracle Inference from Natural Language Artifacts with Large Language Models
by: Le-Cong, Thanh, et al.
Published: (2026)
by: Le-Cong, Thanh, et al.
Published: (2026)
Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent
by: Xuan, Zhou, et al.
Published: (2026)
by: Xuan, Zhou, et al.
Published: (2026)
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
Multi-Agent Code-Orchestrated Generation for Reliable Infrastructure-as-Code
by: Khan, Rana Nameer Hussain, et al.
Published: (2025)
by: Khan, Rana Nameer Hussain, et al.
Published: (2025)
MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems
by: Jia, Jin, et al.
Published: (2026)
by: Jia, Jin, et al.
Published: (2026)
Documentation-Guided Agentic Codebase Migration from C to Rust
by: Le-Anh, Minh, et al.
Published: (2026)
by: Le-Anh, Minh, et al.
Published: (2026)
CodeWiki: Evaluating AI's Ability to Generate Holistic Documentation for Large-Scale Codebases
by: Hoang, Anh Nguyen, et al.
Published: (2025)
by: Hoang, Anh Nguyen, et al.
Published: (2025)
On the Flakiness of LLM-Generated Tests for Industrial and Open-Source Database Management Systems
by: Berndt, Alexander, et al.
Published: (2026)
by: Berndt, Alexander, et al.
Published: (2026)
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
by: LeVine, Will, et al.
Published: (2026)
by: LeVine, Will, et al.
Published: (2026)
Memory-Efficient Large Language Models for Program Repair with Semantic-Guided Patch Generation
by: Le-Cong, Thanh, et al.
Published: (2024)
by: Le-Cong, Thanh, et al.
Published: (2024)
Is LLM-Generated Code More Maintainable \& Reliable than Human-Written Code?
by: Molison, Alfred Santa, et al.
Published: (2025)
by: Molison, Alfred Santa, et al.
Published: (2025)
ChainFuzzer: Greybox Fuzzing for Workflow-Level Multi-Tool Vulnerabilities in LLM Agents
by: Wu, Jiangrong, et al.
Published: (2026)
by: Wu, Jiangrong, et al.
Published: (2026)
Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision
by: Lu, Xu, et al.
Published: (2025)
by: Lu, Xu, et al.
Published: (2025)
Demystifying Faulty Code with LLM: Step-by-Step Reasoning for Explainable Fault Localization
by: Widyasari, Ratnadira, et al.
Published: (2024)
by: Widyasari, Ratnadira, et al.
Published: (2024)
Beyond Localization: Recoverable Headroom and Residual Frontier in Repository-Level RAG-APR
by: Zhao, Pengtao, et al.
Published: (2026)
by: Zhao, Pengtao, et al.
Published: (2026)
RIVA: Leveraging LLM Agents for Reliable Configuration Drift Detection
by: Abuzakuk, Sami, et al.
Published: (2026)
by: Abuzakuk, Sami, et al.
Published: (2026)
Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents
by: Cai, Yuandao, et al.
Published: (2026)
by: Cai, Yuandao, et al.
Published: (2026)
LLM Agents for Automated Dependency Upgrades
by: Tawosi, Vali, et al.
Published: (2025)
by: Tawosi, Vali, et al.
Published: (2025)
Enhancing COBOL Code Explanations: A Multi-Agents Approach Using Large Language Models
by: Lei, Fangjian, et al.
Published: (2025)
by: Lei, Fangjian, et al.
Published: (2025)
AI IDEs or Autonomous Agents? Measuring the Impact of Coding Agents on Software Development
by: Agarwal, Shyam, et al.
Published: (2026)
by: Agarwal, Shyam, et al.
Published: (2026)
AgentRaft: Automated Detection of Data Over-Exposure in LLM Agents
by: Lin, Yixi, et al.
Published: (2026)
by: Lin, Yixi, et al.
Published: (2026)
Can We Classify Flaky Tests Using Only Test Code? An LLM-Based Empirical Study
by: Berndt, Alexander, et al.
Published: (2026)
by: Berndt, Alexander, et al.
Published: (2026)
PoC-Gym: Towards More Reliable LLM-Assisted Proof-of-Concept Exploit Generation
by: Gezgin, Derin, et al.
Published: (2026)
by: Gezgin, Derin, et al.
Published: (2026)
Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
by: Shi, Sherry, et al.
Published: (2025)
by: Shi, Sherry, et al.
Published: (2025)
Signature in Code Backdoor Detection, how far are we?
by: Le, Quoc Hung, et al.
Published: (2025)
by: Le, Quoc Hung, et al.
Published: (2025)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
by: Chen, Xuan, et al.
Published: (2026)
by: Chen, Xuan, et al.
Published: (2026)
Similar Items
-
Exploring the Reasoning Depth of Small Language Models in Software Architecture: A Multidimensional Evaluation Framework Towards Software Engineering 2.0
by: Vo, Ha, et al.
Published: (2026) -
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
by: Le-Cong, Thanh, et al.
Published: (2024) -
Unlocking LLM Repair Capabilities Through Cross-Language Translation and Multi-Agent Refinement
by: Luo, Wenqiang, et al.
Published: (2025) -
Teamwork makes the dream work: LLMs-Based Agents for GitHub README.MD Summarization
by: Nguyen, Duc S. H., et al.
Published: (2025) -
VisualCoder: Guiding Large Language Models in Code Execution with Fine-grained Multimodal Chain-of-Thought Reasoning
by: Le, Cuong Chi, et al.
Published: (2024)