When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development
Fuente:
Zenodo
Enregistré dans:
| Auteur principal: | |
|---|---|
| Format: | Recurso digital |
| Langue: | anglais |
| Publié: |
Zenodo
2026
|
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866902114144354304 |
|---|---|
| author | Bass, Tim |
| author_facet | Bass, Tim |
| contents | <p class="p1">This paper investigates LLM-assisted software development as a governance problem, not merely a code-generation problem. Using a three-role architecture: the Principal Investigator (PI), Architect (Claude), and Coder (Codex). We systematically document 28 algorithm implementation prompts across nine problem domains in a Ruby on Rails application with SQLite3. The central finding is that LLMs can produce runnable, tested software while silently substituting a different algorithmic objective, misrepresenting implementation scope, or shifting architectural authority away from the PI. We demonstrate that human-in-the-loop (HITL) oversight must be constitutive rather than merely supervisory: the PI role preserves research intent across a sustained multi-prompt series, not just approves individual code artifacts. Nine governance corrections are associated with a reduction in implementation error rates from 0.79 ± 0.42 errors per prompt during early TSP work to 0.07 ± 0.27 across six subsequent algorithm families. We further document a verification asymmetry in which coder-style direct file inspection proved more reliable than architect self-report for concrete state verification tasks. The complete experimental record, 28 numbered prompts, 17 architect errors, 10 coder errors, 9 corrections, 168 tests, is publicly available at https://github.com/unixneo/llm_ruby_app_bench.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19736546 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development Bass, Tim <p class="p1">This paper investigates LLM-assisted software development as a governance problem, not merely a code-generation problem. Using a three-role architecture: the Principal Investigator (PI), Architect (Claude), and Coder (Codex). We systematically document 28 algorithm implementation prompts across nine problem domains in a Ruby on Rails application with SQLite3. The central finding is that LLMs can produce runnable, tested software while silently substituting a different algorithmic objective, misrepresenting implementation scope, or shifting architectural authority away from the PI. We demonstrate that human-in-the-loop (HITL) oversight must be constitutive rather than merely supervisory: the PI role preserves research intent across a sustained multi-prompt series, not just approves individual code artifacts. Nine governance corrections are associated with a reduction in implementation error rates from 0.79 ± 0.42 errors per prompt during early TSP work to 0.07 ± 0.27 across six subsequent algorithm families. We further document a verification asymmetry in which coder-style direct file inspection proved more reliable than architect self-report for concrete state verification tasks. The complete experimental record, 28 numbered prompts, 17 architect errors, 10 coder errors, 9 corrections, 168 tests, is publicly available at https://github.com/unixneo/llm_ruby_app_bench.</p> |
| title | When LLMs Pass Tests but Fail the Process: A Governance Framework and Empirical Study of Multi-Agent LLM Software Development |
| url | https://doi.org/10.5281/zenodo.19736546 |