When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Bass, Tim
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2026
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866902112736116736
author Bass, Tim
author_facet Bass, Tim
contents <p class="p1">LLM-assisted software development is fundamentally a governance problem, not merely a code-generation problem. The most serious failure mode is not code that crashes or tests that fail: it is code that runs correctly while answering a different question than the developer intended to ask. This class of failure, which we term goal substitution, is definitionally invisible to automated testing and can only be detected through constitutive human-in-the-loop (HITL) oversight: PI-level approval required before consequential decisions are made, not supervisory review of outputs after the fact. We report a systematic empirical study using a three-role architecture (Principal Investigator (PI), Architect (Claude), and Coder (Codex)) implementing 28 algorithm prompts across nine problem domains in a Ruby on Rails application with SQLite3. We document 30 classified errors using a formal five-type taxonomy (goal substitution, scope limitation, specification error, verification failure, process violation), develop nine governance corrections in response, and observe a reduction in implementation error rates from 0.79 ± 0.42 errors per prompt during early TSP work to 0.33 ± 0.47 across nine subsequent algorithm families. We further document a verification asymmetry in which coder-style direct file inspection proved more reliable than architect self-report for concrete state verification tasks. The complete experimental record, 28 numbered prompts, 17 architect errors, 13 coder errors, 9 corrections, 168 tests, is publicly available at https://github.com/unixneo/llm_ruby_app_bench.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_20364827
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development
Bass, Tim
<p class="p1">LLM-assisted software development is fundamentally a governance problem, not merely a code-generation problem. The most serious failure mode is not code that crashes or tests that fail: it is code that runs correctly while answering a different question than the developer intended to ask. This class of failure, which we term goal substitution, is definitionally invisible to automated testing and can only be detected through constitutive human-in-the-loop (HITL) oversight: PI-level approval required before consequential decisions are made, not supervisory review of outputs after the fact. We report a systematic empirical study using a three-role architecture (Principal Investigator (PI), Architect (Claude), and Coder (Codex)) implementing 28 algorithm prompts across nine problem domains in a Ruby on Rails application with SQLite3. We document 30 classified errors using a formal five-type taxonomy (goal substitution, scope limitation, specification error, verification failure, process violation), develop nine governance corrections in response, and observe a reduction in implementation error rates from 0.79 ± 0.42 errors per prompt during early TSP work to 0.33 ± 0.47 across nine subsequent algorithm families. We further document a verification asymmetry in which coder-style direct file inspection proved more reliable than architect self-report for concrete state verification tasks. The complete experimental record, 28 numbered prompts, 17 architect errors, 13 coder errors, 9 corrections, 168 tests, is publicly available at https://github.com/unixneo/llm_ruby_app_bench.</p>
title When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development
url https://doi.org/10.5281/zenodo.20364827