When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | Bass, Tim |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2026
|
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development
von: Bass, Tim
Veröffentlicht: (2026)
von: Bass, Tim
Veröffentlicht: (2026)
Testing the Untestable? An Empirical Study on the Testing Process of LLM-Powered Software Systems
von: Magalhaes, Cleyton, et al.
Veröffentlicht: (2025)
von: Magalhaes, Cleyton, et al.
Veröffentlicht: (2025)
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
von: Rehman, Mohammad Abdul, et al.
Veröffentlicht: (2025)
von: Rehman, Mohammad Abdul, et al.
Veröffentlicht: (2025)
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling
von: Islam, Niful, et al.
Veröffentlicht: (2026)
von: Islam, Niful, et al.
Veröffentlicht: (2026)
When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
von: Li, Xiaoxiao
Veröffentlicht: (2026)
von: Li, Xiaoxiao
Veröffentlicht: (2026)
The Z₂₄ Challenge: Three Hardened Pass/Fail Superconducting Tests for a Crystalline Axiverse Governance Rule
von: Diogenes
Veröffentlicht: (2026)
von: Diogenes
Veröffentlicht: (2026)
Beyond Pass/Fail: The Story of Learning-Based Testing
von: Rahman, Sheikh Md. Mushfiqur, et al.
Veröffentlicht: (2025)
von: Rahman, Sheikh Md. Mushfiqur, et al.
Veröffentlicht: (2025)
When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems
von: Zhang, Lingxi, et al.
Veröffentlicht: (2026)
von: Zhang, Lingxi, et al.
Veröffentlicht: (2026)
LLMs are Imperfect, Then What? An Empirical Study on LLM Failures in Software Engineering
von: Anonymous, Anonymous
Veröffentlicht: (2024)
von: Anonymous, Anonymous
Veröffentlicht: (2024)
LLMs are Imperfect, Then What? An Empirical Study on LLM Failures in Software Engineering
von: Anonymous, Anonymous
Veröffentlicht: (2024)
von: Anonymous, Anonymous
Veröffentlicht: (2024)
LLMs are Imperfect, Then What? An Empirical Study on LLM Failures in Software Engineering
von: Anonymous, Anonymous
Veröffentlicht: (2025)
von: Anonymous, Anonymous
Veröffentlicht: (2025)
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
A Framework for Using LLMs for Repository Mining Studies in Empirical Software Engineering
von: de Martino, Vincenzo, et al.
Veröffentlicht: (2024)
von: de Martino, Vincenzo, et al.
Veröffentlicht: (2024)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
Why Do Multi-Agent LLM Systems Fail?
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
Multi-Agent LLM Committees for Autonomous Software Beta Testing
von: Karanam, Sumanth Bharadwaj Hachalli, et al.
Veröffentlicht: (2025)
von: Karanam, Sumanth Bharadwaj Hachalli, et al.
Veröffentlicht: (2025)
An Empirical Study of Agent Developer Practices in AI Agent Frameworks
von: Wang, Yanlin, et al.
Veröffentlicht: (2025)
von: Wang, Yanlin, et al.
Veröffentlicht: (2025)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
von: Ehsani, Ramtin, et al.
Veröffentlicht: (2026)
von: Ehsani, Ramtin, et al.
Veröffentlicht: (2026)
When KV Cache Reuse Fails in Multi-Agent Systems: Cross-Candidate Interaction is Crucial for LLM Judges
von: Liang, Sichu, et al.
Veröffentlicht: (2026)
von: Liang, Sichu, et al.
Veröffentlicht: (2026)
An Empirical Study of Bugs in Modern LLM Agent Frameworks
von: Zhu, Xinxue, et al.
Veröffentlicht: (2026)
von: Zhu, Xinxue, et al.
Veröffentlicht: (2026)
Replication kit for: Can LLMs Make Software Testing Greener? An Empirical Study on JUnit Test Energy Reengineering
von: Anonymous
Veröffentlicht: (2025)
von: Anonymous
Veröffentlicht: (2025)
M2-PALE: A Framework for Explaining Multi-Agent MCTS--Minimax Hybrids via Process Mining and LLMs
von: Qian, Yiyu, et al.
Veröffentlicht: (2026)
von: Qian, Yiyu, et al.
Veröffentlicht: (2026)
The Silent Scientist: When Software Research Fails to Reach Its Audience
von: Wyrich, Marvin, et al.
Veröffentlicht: (2025)
von: Wyrich, Marvin, et al.
Veröffentlicht: (2025)
When Rituals Fail: Rationalization, Bayesianism, and Predictive Processing
von: Ze Hong
Veröffentlicht: (2025)
von: Ze Hong
Veröffentlicht: (2025)
The Neimheadh Framework: Geometric Unity and Toroidal Field Interactions
von: Bass, Joseph
Veröffentlicht: (2025)
von: Bass, Joseph
Veröffentlicht: (2025)
An Empirical Study on the Potential of LLMs in Automated Software Refactoring
von: Liu, Bo, et al.
Veröffentlicht: (2024)
von: Liu, Bo, et al.
Veröffentlicht: (2024)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
A Methodological Analysis of Empirical Studies in Quantum Software Testing
von: Li, Yuechen, et al.
Veröffentlicht: (2026)
von: Li, Yuechen, et al.
Veröffentlicht: (2026)
Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development
von: Shen, Ming, et al.
Veröffentlicht: (2025)
von: Shen, Ming, et al.
Veröffentlicht: (2025)
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
von: Li, Xiaomin, et al.
Veröffentlicht: (2025)
von: Li, Xiaomin, et al.
Veröffentlicht: (2025)
When Reasoning Fails: Evaluating 'Thinking' LLMs for Stock Prediction
von: Sodha, Rakeshkumar H
Veröffentlicht: (2025)
von: Sodha, Rakeshkumar H
Veröffentlicht: (2025)
Lost in Transmission: When and Why LLMs Fail to Reason Globally
von: Schnabel, Tobias, et al.
Veröffentlicht: (2025)
von: Schnabel, Tobias, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?
von: Mayor-Rocher, Marina, et al.
Veröffentlicht: (2024)
von: Mayor-Rocher, Marina, et al.
Veröffentlicht: (2024)
When Explanations Lie: Why Many Modified BP Attributions Fail
von: Sixt, Leon, et al.
Veröffentlicht: (2019)
von: Sixt, Leon, et al.
Veröffentlicht: (2019)
When Data Protection Fails to Protect: Law, Power, and Postcolonial Governance in Bangladesh
von: Saha, Pratyasha, et al.
Veröffentlicht: (2026)
von: Saha, Pratyasha, et al.
Veröffentlicht: (2026)
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
von: Paglieri, Davide, et al.
Veröffentlicht: (2025)
von: Paglieri, Davide, et al.
Veröffentlicht: (2025)
A replication package for " An Empirical Study of Bugs in LLM-based Agent Frameworks"
von: Batole, Fraol
Veröffentlicht: (2026)
von: Batole, Fraol
Veröffentlicht: (2026)
When Identity Overrides Incentives: Representational Choices as Governance Decisions in Multi-Agent LLM Systems
von: Manoranjan, Viswonathan, et al.
Veröffentlicht: (2026)
von: Manoranjan, Viswonathan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development
von: Bass, Tim
Veröffentlicht: (2026) -
Testing the Untestable? An Empirical Study on the Testing Process of LLM-Powered Software Systems
von: Magalhaes, Cleyton, et al.
Veröffentlicht: (2025) -
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
von: Rehman, Mohammad Abdul, et al.
Veröffentlicht: (2025) -
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025) -
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
von: Huang, Donghao, et al.
Veröffentlicht: (2026)