When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2026
|
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866902112736116736 |
|---|---|
| author | Bass, Tim |
| author_facet | Bass, Tim |
| contents | <p class="p1">LLM-assisted software development is fundamentally a governance problem, not merely a code-generation problem. The most serious failure mode is not code that crashes or tests that fail: it is code that runs correctly while answering a different question than the developer intended to ask. This class of failure, which we term goal substitution, is definitionally invisible to automated testing and can only be detected through constitutive human-in-the-loop (HITL) oversight: PI-level approval required before consequential decisions are made, not supervisory review of outputs after the fact. We report a systematic empirical study using a three-role architecture (Principal Investigator (PI), Architect (Claude), and Coder (Codex)) implementing 28 algorithm prompts across nine problem domains in a Ruby on Rails application with SQLite3. We document 30 classified errors using a formal five-type taxonomy (goal substitution, scope limitation, specification error, verification failure, process violation), develop nine governance corrections in response, and observe a reduction in implementation error rates from 0.79 ± 0.42 errors per prompt during early TSP work to 0.33 ± 0.47 across nine subsequent algorithm families. We further document a verification asymmetry in which coder-style direct file inspection proved more reliable than architect self-report for concrete state verification tasks. The complete experimental record, 28 numbered prompts, 17 architect errors, 13 coder errors, 9 corrections, 168 tests, is publicly available at https://github.com/unixneo/llm_ruby_app_bench.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_20364827 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development Bass, Tim <p class="p1">LLM-assisted software development is fundamentally a governance problem, not merely a code-generation problem. The most serious failure mode is not code that crashes or tests that fail: it is code that runs correctly while answering a different question than the developer intended to ask. This class of failure, which we term goal substitution, is definitionally invisible to automated testing and can only be detected through constitutive human-in-the-loop (HITL) oversight: PI-level approval required before consequential decisions are made, not supervisory review of outputs after the fact. We report a systematic empirical study using a three-role architecture (Principal Investigator (PI), Architect (Claude), and Coder (Codex)) implementing 28 algorithm prompts across nine problem domains in a Ruby on Rails application with SQLite3. We document 30 classified errors using a formal five-type taxonomy (goal substitution, scope limitation, specification error, verification failure, process violation), develop nine governance corrections in response, and observe a reduction in implementation error rates from 0.79 ± 0.42 errors per prompt during early TSP work to 0.33 ± 0.47 across nine subsequent algorithm families. We further document a verification asymmetry in which coder-style direct file inspection proved more reliable than architect self-report for concrete state verification tasks. The complete experimental record, 28 numbered prompts, 17 architect errors, 13 coder errors, 9 corrections, 168 tests, is publicly available at https://github.com/unixneo/llm_ruby_app_bench.</p> |
| title | When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development |
| url | https://doi.org/10.5281/zenodo.20364827 |