How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
Fuente:
arXiv
Saved in:
| Main Author: | Roig, JV |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs for Qualitative Data Analysis Fail on Security-specificComments in Human Experiments
by: Camporese, Maria, et al.
Published: (2026)
by: Camporese, Maria, et al.
Published: (2026)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
by: Ehsani, Ramtin, et al.
Published: (2026)
by: Ehsani, Ramtin, et al.
Published: (2026)
Demystifying the Lifecycle of Failures in Platform-Orchestrated Agentic Workflows
by: Ma, Xuyan, et al.
Published: (2025)
by: Ma, Xuyan, et al.
Published: (2025)
Willful Disobedience: Automatically Detecting Failures in Agentic Traces
by: Sharma, Reshabh K, et al.
Published: (2026)
by: Sharma, Reshabh K, et al.
Published: (2026)
Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations
by: Honarvar, Shahin, et al.
Published: (2026)
by: Honarvar, Shahin, et al.
Published: (2026)
How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms
by: Roig, JV
Published: (2026)
by: Roig, JV
Published: (2026)
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?
by: Garg, Spandan, et al.
Published: (2026)
by: Garg, Spandan, et al.
Published: (2026)
ROSBag MCP Server: Analyzing Robot Data with LLMs for Agentic Embodied AI Applications
by: Fu, Lei, et al.
Published: (2025)
by: Fu, Lei, et al.
Published: (2025)
PEFA-AI: Advancing Open-source LLMs for RTL generation using Progressive Error Feedback Agentic-AI
by: Narayanan, Athma, et al.
Published: (2025)
by: Narayanan, Athma, et al.
Published: (2025)
Perceptual Self-Reflection in Agentic Physics Simulation Code Generation
by: Shende, Prashant, et al.
Published: (2026)
by: Shende, Prashant, et al.
Published: (2026)
Beyond the 'Diff': Addressing Agentic Entropy in Agentic Software Development
by: Casserini, Matteo, et al.
Published: (2026)
by: Casserini, Matteo, et al.
Published: (2026)
From Inductive to Deductive: LLMs-Based Qualitative Data Analysis in Requirements Engineering
by: Shah, Syed Tauhid Ullah, et al.
Published: (2025)
by: Shah, Syed Tauhid Ullah, et al.
Published: (2025)
Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?
by: Pan, Zhenyu, et al.
Published: (2024)
by: Pan, Zhenyu, et al.
Published: (2024)
Agentic Scientific Simulation: Execution-Grounded Model Construction and Reconstruction
by: Lie, Knut-Andreas, et al.
Published: (2026)
by: Lie, Knut-Andreas, et al.
Published: (2026)
On Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset
by: Bhat, Vishvesh, et al.
Published: (2025)
by: Bhat, Vishvesh, et al.
Published: (2025)
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
by: Feng, Dylan, et al.
Published: (2026)
by: Feng, Dylan, et al.
Published: (2026)
Do Code LLMs Understand Design Patterns?
by: Pan, Zhenyu, et al.
Published: (2025)
by: Pan, Zhenyu, et al.
Published: (2025)
Pragmos: A Process Agentic Modeling System
by: Hernández-Ávalos, Pedro-Aarón, et al.
Published: (2026)
by: Hernández-Ávalos, Pedro-Aarón, et al.
Published: (2026)
DeepCode: Open Agentic Coding
by: Li, Zongwei, et al.
Published: (2025)
by: Li, Zongwei, et al.
Published: (2025)
Agentic Business Process Management Systems
by: Dumas, Marlon, et al.
Published: (2026)
by: Dumas, Marlon, et al.
Published: (2026)
Agentic Harness for Real-World Compilers
by: Zheng, Yingwei, et al.
Published: (2026)
by: Zheng, Yingwei, et al.
Published: (2026)
Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
by: Jin, Haolin, et al.
Published: (2025)
by: Jin, Haolin, et al.
Published: (2025)
Agentic AI Software Engineers: Programming with Trust
by: Roychoudhury, Abhik, et al.
Published: (2025)
by: Roychoudhury, Abhik, et al.
Published: (2025)
Agentic Coding Needs Proactivity, Not Just Autonomy
by: Bui, Nghi D. Q., et al.
Published: (2026)
by: Bui, Nghi D. Q., et al.
Published: (2026)
Agentic Frameworks for Reasoning Tasks: An Empirical Study
by: Rasheed, Zeeshan, et al.
Published: (2026)
by: Rasheed, Zeeshan, et al.
Published: (2026)
ScenEval: A Benchmark for Scenario-Based Evaluation of Code Generation
by: Paul, Debalina Ghosh, et al.
Published: (2024)
by: Paul, Debalina Ghosh, et al.
Published: (2024)
Text2Scenario: Text-Driven Scenario Generation for Autonomous Driving Test
by: Cai, Xuan, et al.
Published: (2025)
by: Cai, Xuan, et al.
Published: (2025)
Process-Centric Analysis of Agentic Software Systems
by: Liu, Shuyang, et al.
Published: (2025)
by: Liu, Shuyang, et al.
Published: (2025)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
by: Gao, Zeyu, et al.
Published: (2025)
by: Gao, Zeyu, et al.
Published: (2025)
ARPaCCino: An Agentic-RAG for Policy as Code Compliance
by: Romeo, Francesco, et al.
Published: (2025)
by: Romeo, Francesco, et al.
Published: (2025)
BabelCoder: Agentic Code Translation with Specification Alignment
by: Rabbi, Fazle, et al.
Published: (2025)
by: Rabbi, Fazle, et al.
Published: (2025)
Runtime-Structured Task Decomposition for Agentic Coding Systems
by: Asthana, Shubhi, et al.
Published: (2026)
by: Asthana, Shubhi, et al.
Published: (2026)
BLAgent: Agentic RAG for File-Level Bug Localization
by: Mamun, Md Afif Al, et al.
Published: (2026)
by: Mamun, Md Afif Al, et al.
Published: (2026)
Quantifying the Expectation-Realisation Gap for Agentic AI Systems
by: Lobentanzer, Sebastian
Published: (2026)
by: Lobentanzer, Sebastian
Published: (2026)
Analysis on LLMs Performance for Code Summarization
by: Akib, Md. Ahnaf, et al.
Published: (2024)
by: Akib, Md. Ahnaf, et al.
Published: (2024)
Agentic Software Issue Resolution with Large Language Models: A Survey
by: Jiang, Zhonghao, et al.
Published: (2025)
by: Jiang, Zhonghao, et al.
Published: (2025)
TriCEGAR: A Trace-Driven Abstraction Mechanism for Agentic AI
by: Koohestani, Roham, et al.
Published: (2026)
by: Koohestani, Roham, et al.
Published: (2026)
MLDebugging: Towards Benchmarking Code Debugging Across Multi-Library Scenarios
by: Huang, Jinyang, et al.
Published: (2025)
by: Huang, Jinyang, et al.
Published: (2025)
SPARC: Scenario Planning and Reasoning for Automated C Unit Test Generation
by: Chowdhury, Jaid Monwar, et al.
Published: (2026)
by: Chowdhury, Jaid Monwar, et al.
Published: (2026)
InCoder-32B: Code Foundation Model for Industrial Scenarios
by: Yang, Jian, et al.
Published: (2026)
by: Yang, Jian, et al.
Published: (2026)
Similar Items
-
LLMs for Qualitative Data Analysis Fail on Security-specificComments in Human Experiments
by: Camporese, Maria, et al.
Published: (2026) -
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
by: Ehsani, Ramtin, et al.
Published: (2026) -
Demystifying the Lifecycle of Failures in Platform-Orchestrated Agentic Workflows
by: Ma, Xuyan, et al.
Published: (2025) -
Willful Disobedience: Automatically Detecting Failures in Agentic Traces
by: Sharma, Reshabh K, et al.
Published: (2026) -
Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations
by: Honarvar, Shahin, et al.
Published: (2026)