Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Wang Bill, Chai, Miaosen, Wang, Shangshang, Liu, Yejia, Bian, Song, Dong, Honghua, Neiswanger, Willie, Jia, Robin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918505165619200
author Zhu, Wang Bill
Chai, Miaosen
Wang, Shangshang
Liu, Yejia
Bian, Song
Dong, Honghua
Neiswanger, Willie
Jia, Robin
author_facet Zhu, Wang Bill
Chai, Miaosen
Wang, Shangshang
Liu, Yejia
Bian, Song
Dong, Honghua
Neiswanger, Willie
Jia, Robin
contents Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but over-edited solutions during debugging. To evaluate how far LLMs are from precise debugging, we introduce the Precise Debugging Benchmark (PDB) framework, which automatically converts any coding dataset into a debugging benchmark with precision-aware evaluation. PDB generates buggy programs by synthesizing verified atomic bugs and composing them into multi-bug programs. We define two novel metrics, edit-level precision and bug-level recall, which measures how many necessary edits are made and how many bugs are resolved. We release two evaluation benchmarks: PDB-Single-Hard on single-line bugs, and PDB-Multi on multi-line bugs. Experiments show that frontier models, such as GPT-5.1-Codex and DeepSeek-V3.2-Thinking, achieve unit-test pass rates above 76% but exhibit precision below 45%, even when explicitly instructed to perform minimal debugging. Finally, we show that iterative and agentic debugging strategies do not substantially improve precision or recall, highlighting the need to rethink post-training pipelines for coding models.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17338
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
Zhu, Wang Bill
Chai, Miaosen
Wang, Shangshang
Liu, Yejia
Bian, Song
Dong, Honghua
Neiswanger, Willie
Jia, Robin
Software Engineering
Computation and Language
Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but over-edited solutions during debugging. To evaluate how far LLMs are from precise debugging, we introduce the Precise Debugging Benchmark (PDB) framework, which automatically converts any coding dataset into a debugging benchmark with precision-aware evaluation. PDB generates buggy programs by synthesizing verified atomic bugs and composing them into multi-bug programs. We define two novel metrics, edit-level precision and bug-level recall, which measures how many necessary edits are made and how many bugs are resolved. We release two evaluation benchmarks: PDB-Single-Hard on single-line bugs, and PDB-Multi on multi-line bugs. Experiments show that frontier models, such as GPT-5.1-Codex and DeepSeek-V3.2-Thinking, achieve unit-test pass rates above 76% but exhibit precision below 45%, even when explicitly instructed to perform minimal debugging. Finally, we show that iterative and agentic debugging strategies do not substantially improve precision or recall, highlighting the need to rethink post-training pipelines for coding models.
title Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
topic Software Engineering
Computation and Language
url https://arxiv.org/abs/2604.17338