When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development
Fuente:
Zenodo
Saved in:
| Main Author: | Bass, Tim |
|---|---|
| Format: | Recurso digital |
| Language: | English |
| Published: |
Zenodo
2026
|
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development
by: Bass, Tim
Published: (2026)
by: Bass, Tim
Published: (2026)
Testing the Untestable? An Empirical Study on the Testing Process of LLM-Powered Software Systems
by: Magalhaes, Cleyton, et al.
Published: (2025)
by: Magalhaes, Cleyton, et al.
Published: (2025)
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
by: Huang, Donghao, et al.
Published: (2026)
by: Huang, Donghao, et al.
Published: (2026)
When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling
by: Islam, Niful, et al.
Published: (2026)
by: Islam, Niful, et al.
Published: (2026)
When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
by: Li, Xiaoxiao
Published: (2026)
by: Li, Xiaoxiao
Published: (2026)
The Z₂₄ Challenge: Three Hardened Pass/Fail Superconducting Tests for a Crystalline Axiverse Governance Rule
by: Diogenes
Published: (2026)
by: Diogenes
Published: (2026)
Beyond Pass/Fail: The Story of Learning-Based Testing
by: Rahman, Sheikh Md. Mushfiqur, et al.
Published: (2025)
by: Rahman, Sheikh Md. Mushfiqur, et al.
Published: (2025)
When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems
by: Zhang, Lingxi, et al.
Published: (2026)
by: Zhang, Lingxi, et al.
Published: (2026)
LLMs are Imperfect, Then What? An Empirical Study on LLM Failures in Software Engineering
by: Anonymous, Anonymous
Published: (2024)
by: Anonymous, Anonymous
Published: (2024)
LLMs are Imperfect, Then What? An Empirical Study on LLM Failures in Software Engineering
by: Anonymous, Anonymous
Published: (2024)
by: Anonymous, Anonymous
Published: (2024)
LLMs are Imperfect, Then What? An Empirical Study on LLM Failures in Software Engineering
by: Anonymous, Anonymous
Published: (2025)
by: Anonymous, Anonymous
Published: (2025)
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
by: Wang, Zehao, et al.
Published: (2026)
by: Wang, Zehao, et al.
Published: (2026)
A Framework for Using LLMs for Repository Mining Studies in Empirical Software Engineering
by: de Martino, Vincenzo, et al.
Published: (2024)
by: de Martino, Vincenzo, et al.
Published: (2024)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
by: Hadeliya, Tsimur, et al.
Published: (2025)
by: Hadeliya, Tsimur, et al.
Published: (2025)
Why Do Multi-Agent LLM Systems Fail?
by: Cemri, Mert, et al.
Published: (2025)
by: Cemri, Mert, et al.
Published: (2025)
Multi-Agent LLM Committees for Autonomous Software Beta Testing
by: Karanam, Sumanth Bharadwaj Hachalli, et al.
Published: (2025)
by: Karanam, Sumanth Bharadwaj Hachalli, et al.
Published: (2025)
An Empirical Study of Agent Developer Practices in AI Agent Frameworks
by: Wang, Yanlin, et al.
Published: (2025)
by: Wang, Yanlin, et al.
Published: (2025)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
by: Ehsani, Ramtin, et al.
Published: (2026)
by: Ehsani, Ramtin, et al.
Published: (2026)
When KV Cache Reuse Fails in Multi-Agent Systems: Cross-Candidate Interaction is Crucial for LLM Judges
by: Liang, Sichu, et al.
Published: (2026)
by: Liang, Sichu, et al.
Published: (2026)
An Empirical Study of Bugs in Modern LLM Agent Frameworks
by: Zhu, Xinxue, et al.
Published: (2026)
by: Zhu, Xinxue, et al.
Published: (2026)
Replication kit for: Can LLMs Make Software Testing Greener? An Empirical Study on JUnit Test Energy Reengineering
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
M2-PALE: A Framework for Explaining Multi-Agent MCTS--Minimax Hybrids via Process Mining and LLMs
by: Qian, Yiyu, et al.
Published: (2026)
by: Qian, Yiyu, et al.
Published: (2026)
The Silent Scientist: When Software Research Fails to Reach Its Audience
by: Wyrich, Marvin, et al.
Published: (2025)
by: Wyrich, Marvin, et al.
Published: (2025)
When Rituals Fail: Rationalization, Bayesianism, and Predictive Processing
by: Ze Hong
Published: (2025)
by: Ze Hong
Published: (2025)
The Neimheadh Framework: Geometric Unity and Toroidal Field Interactions
by: Bass, Joseph
Published: (2025)
by: Bass, Joseph
Published: (2025)
An Empirical Study on the Potential of LLMs in Automated Software Refactoring
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
A Methodological Analysis of Empirical Studies in Quantum Software Testing
by: Li, Yuechen, et al.
Published: (2026)
by: Li, Yuechen, et al.
Published: (2026)
Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development
by: Shen, Ming, et al.
Published: (2025)
by: Shen, Ming, et al.
Published: (2025)
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
by: Li, Xiaomin, et al.
Published: (2025)
by: Li, Xiaomin, et al.
Published: (2025)
When Reasoning Fails: Evaluating 'Thinking' LLMs for Stock Prediction
by: Sodha, Rakeshkumar H
Published: (2025)
by: Sodha, Rakeshkumar H
Published: (2025)
Lost in Transmission: When and Why LLMs Fail to Reason Globally
by: Schnabel, Tobias, et al.
Published: (2025)
by: Schnabel, Tobias, et al.
Published: (2025)
Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?
by: Mayor-Rocher, Marina, et al.
Published: (2024)
by: Mayor-Rocher, Marina, et al.
Published: (2024)
When Explanations Lie: Why Many Modified BP Attributions Fail
by: Sixt, Leon, et al.
Published: (2019)
by: Sixt, Leon, et al.
Published: (2019)
When Data Protection Fails to Protect: Law, Power, and Postcolonial Governance in Bangladesh
by: Saha, Pratyasha, et al.
Published: (2026)
by: Saha, Pratyasha, et al.
Published: (2026)
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
by: Paglieri, Davide, et al.
Published: (2025)
by: Paglieri, Davide, et al.
Published: (2025)
A replication package for " An Empirical Study of Bugs in LLM-based Agent Frameworks"
by: Batole, Fraol
Published: (2026)
by: Batole, Fraol
Published: (2026)
When Identity Overrides Incentives: Representational Choices as Governance Decisions in Multi-Agent LLM Systems
by: Manoranjan, Viswonathan, et al.
Published: (2026)
by: Manoranjan, Viswonathan, et al.
Published: (2026)
Similar Items
-
When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development
by: Bass, Tim
Published: (2026) -
Testing the Untestable? An Empirical Study on the Testing Process of LLM-Powered Software Systems
by: Magalhaes, Cleyton, et al.
Published: (2025) -
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
by: Rehman, Mohammad Abdul, et al.
Published: (2025) -
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025) -
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
by: Huang, Donghao, et al.
Published: (2026)