CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zaoyu, Dai, Jianbo, Zhu, Boyu, Wang, Jingdong, Wang, Huiming, Xu, Xin, Yuan, Haoyang, Guo, Zhijiang, Wu, Xiao-Ming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914470954008576
author Chen, Zaoyu
Dai, Jianbo
Zhu, Boyu
Wang, Jingdong
Wang, Huiming
Xu, Xin
Yuan, Haoyang
Guo, Zhijiang
Wu, Xiao-Ming
author_facet Chen, Zaoyu
Dai, Jianbo
Zhu, Boyu
Wang, Jingdong
Wang, Huiming
Xu, Xin
Yuan, Haoyang
Guo, Zhijiang
Wu, Xiao-Ming
contents Large language models (LLMs) can generate code from natural language, but the extent to which they capture intended program behavior remains unclear. Executable behavioral specifications, defined via preconditions and postconditions, provide a concrete means to assess such understanding. However, existing work on specification generation is constrained in evaluation methodology, task settings, and specification expressiveness. We introduce CodeSpecBench, a benchmark for executable behavioral specification generation under an execution-based evaluation protocol. CodeSpecBench supports both function-level and repository-level tasks and encodes specifications as executable Python functions. Constructed from diverse real-world codebases, it enables a realistic assessment of both correctness (accepting valid behaviors) and completeness (rejecting invalid behaviors). Evaluating 15 state-of-the-art LLMs on CodeSpecBench, we observe a sharp performance degradation on repository-level tasks, where the best model attains only a 20.2% pass rate. We further find that specification generation is substantially more challenging than code generation, indicating that strong coding performance does not necessarily reflect deep understanding of intended program semantics. Our data and code are available at https://github.com/SparksofAGI/CodeSpecBench.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12268
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
Chen, Zaoyu
Dai, Jianbo
Zhu, Boyu
Wang, Jingdong
Wang, Huiming
Xu, Xin
Yuan, Haoyang
Guo, Zhijiang
Wu, Xiao-Ming
Software Engineering
Computation and Language
Large language models (LLMs) can generate code from natural language, but the extent to which they capture intended program behavior remains unclear. Executable behavioral specifications, defined via preconditions and postconditions, provide a concrete means to assess such understanding. However, existing work on specification generation is constrained in evaluation methodology, task settings, and specification expressiveness. We introduce CodeSpecBench, a benchmark for executable behavioral specification generation under an execution-based evaluation protocol. CodeSpecBench supports both function-level and repository-level tasks and encodes specifications as executable Python functions. Constructed from diverse real-world codebases, it enables a realistic assessment of both correctness (accepting valid behaviors) and completeness (rejecting invalid behaviors). Evaluating 15 state-of-the-art LLMs on CodeSpecBench, we observe a sharp performance degradation on repository-level tasks, where the best model attains only a 20.2% pass rate. We further find that specification generation is substantially more challenging than code generation, indicating that strong coding performance does not necessarily reflect deep understanding of intended program semantics. Our data and code are available at https://github.com/SparksofAGI/CodeSpecBench.
title CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
topic Software Engineering
Computation and Language
url https://arxiv.org/abs/2604.12268