ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Jiawei, Jia, Jie, Liu, Chenbo, Xue, Chaoyi, Song, Yapeng, Yang, Xikai, Sun, Dong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917535950045184
author He, Jiawei
Jia, Jie
Liu, Chenbo
Xue, Chaoyi
Song, Yapeng
Yang, Xikai
Sun, Dong
author_facet He, Jiawei
Jia, Jie
Liu, Chenbo
Xue, Chaoyi
Song, Yapeng
Yang, Xikai
Sun, Dong
contents Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limited visibility and often miss defects that arise during execution. We present ProcCtrlBench, a benchmark for execution-process evaluation in LLM coding agents. ProcCtrlBench organizes recurrent execution defects into a reusable ontology covering 11 defect types in 4 categories, and evaluates agent trajectories through standardized process evidence rather than final outcomes alone. To support comparison across heterogeneous agents, ProcCtrlBench standardizes raw logs into a unified trajectory representation and reports calibrated scorecards over process-level findings. In addition, ProcCtrlBench uses control preservation as a way to quantify execution-process quality, capturing whether execution remains interpretable, interruptible, correctable, reversible, and able to hand back authority when needed. We evaluate ProcCtrlBench on 200 cases sampled from three benchmarks: AndroidBench, TerminalBench, and SWE-bench-Verified. Results show that ProcCtrlBench can be instantiated with useful reliability, provides more stable semantics than direct thresholding, and reveals meaningful differences in execution quality that are often overlooked by conventional outcome-based evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20251
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
He, Jiawei
Jia, Jie
Liu, Chenbo
Xue, Chaoyi
Song, Yapeng
Yang, Xikai
Sun, Dong
Software Engineering
Artificial Intelligence
Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limited visibility and often miss defects that arise during execution. We present ProcCtrlBench, a benchmark for execution-process evaluation in LLM coding agents. ProcCtrlBench organizes recurrent execution defects into a reusable ontology covering 11 defect types in 4 categories, and evaluates agent trajectories through standardized process evidence rather than final outcomes alone. To support comparison across heterogeneous agents, ProcCtrlBench standardizes raw logs into a unified trajectory representation and reports calibrated scorecards over process-level findings. In addition, ProcCtrlBench uses control preservation as a way to quantify execution-process quality, capturing whether execution remains interpretable, interruptible, correctable, reversible, and able to hand back authority when needed. We evaluate ProcCtrlBench on 200 cases sampled from three benchmarks: AndroidBench, TerminalBench, and SWE-bench-Verified. Results show that ProcCtrlBench can be instantiated with useful reliability, provides more stable semantics than direct thresholding, and reveals meaningful differences in execution quality that are often overlooked by conventional outcome-based evaluation.
title ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2605.20251