Guardado en:
Detalles Bibliográficos
Autores principales: Orlanski, Gabriel, Roy, Devjeet, Yun, Alexander, Shin, Changho, Gu, Alex, Ge, Albert, Adila, Dyah, Roberts, Nicholas, Sala, Frederic, Albarghouthi, Aws
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:https://arxiv.org/abs/2603.24755
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909024807550976
author Orlanski, Gabriel
Roy, Devjeet
Yun, Alexander
Shin, Changho
Gu, Alex
Ge, Albert
Adila, Dyah
Roberts, Nicholas
Sala, Frederic
Albarghouthi, Aws
author_facet Orlanski, Gabriel
Roy, Devjeet
Yun, Alexander
Shin, Changho
Gu, Alex
Ge, Albert
Adila, Dyah
Roberts, Nicholas
Sala, Frederic
Albarghouthi, Aws
contents Software development is iterative, yet agentic coding benchmarks hide design issues through their single-shot setup. Recent iterative benchmarks attempt to remedy this but heavily constrain an agent's design decision space, making it impossible to faithfully measure how their decisions shape future extensions. We introduce SlopCodeBench, a benchmark of 36 problems and 196 checkpoints where agents repeatedly extend their own solutions. Unlike prior iterative benchmarks, our evolving specifications demand architectural decisions but leave internal structure to the agent. We measure two forms of degradation: structural erosion (concentrated complexity) and verbosity (redundant code). Evaluating 15 coding agents across open and closed models, we find that no agent fully solves any problem end-to-end, and the best agent passes 14.8% of checkpoints. Quality degrades across checkpoints, with structural erosion rising in 77% of trajectories and verbosity in 75.5%. Compared to 473 open-source Python repositories, agent code is 2.3x more verbose and 2.0x more eroded, and the human repositories degrade less often and by smaller margins across their git histories. Explicit quality guidance reduces initial verbosity and erosion by up to a third, without affecting degradation rates. SlopCodeBench provides the first measurement of code degradation under iterative extension, revealing that agents pass checkpoints while producing code that erodes and bloats with each turn.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24755
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
Orlanski, Gabriel
Roy, Devjeet
Yun, Alexander
Shin, Changho
Gu, Alex
Ge, Albert
Adila, Dyah
Roberts, Nicholas
Sala, Frederic
Albarghouthi, Aws
Software Engineering
Artificial Intelligence
Computation and Language
Software development is iterative, yet agentic coding benchmarks hide design issues through their single-shot setup. Recent iterative benchmarks attempt to remedy this but heavily constrain an agent's design decision space, making it impossible to faithfully measure how their decisions shape future extensions. We introduce SlopCodeBench, a benchmark of 36 problems and 196 checkpoints where agents repeatedly extend their own solutions. Unlike prior iterative benchmarks, our evolving specifications demand architectural decisions but leave internal structure to the agent. We measure two forms of degradation: structural erosion (concentrated complexity) and verbosity (redundant code). Evaluating 15 coding agents across open and closed models, we find that no agent fully solves any problem end-to-end, and the best agent passes 14.8% of checkpoints. Quality degrades across checkpoints, with structural erosion rising in 77% of trajectories and verbosity in 75.5%. Compared to 473 open-source Python repositories, agent code is 2.3x more verbose and 2.0x more eroded, and the human repositories degrade less often and by smaller margins across their git histories. Explicit quality guidance reduces initial verbosity and erosion by up to a third, without affecting degradation rates. SlopCodeBench provides the first measurement of code degradation under iterative extension, revealing that agents pass checkpoints while producing code that erodes and bloats with each turn.
title SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.24755