CIFE: Code Instruction-Following Evaluation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gunnu, Sravani, Guttula, Shanmukha, Patel, Hima
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909970807652352
author Gunnu, Sravani
Guttula, Shanmukha
Patel, Hima
author_facet Gunnu, Sravani
Guttula, Shanmukha
Patel, Hima
contents Large Language Models (LLMs) are increasingly applied to real-world code generation, where functional correctness alone is insufficient for reliable deployment, developers also expect adherence to explicit requirements for robustness, formatting, and security. Existing benchmarks primarily assess correctness through test-case execution, offering limited insight into how reliably models follow such constraints. We introduce a benchmark of 1,000 Python tasks, each paired with an average of 7 developer-specified constraints spanning 13 categories. Constraints are curated through a four-stage human-LLM pipeline to ensure they are atomic, relevant, and objective. We evaluate 14 open- and closed-source models using complementary adherence metrics and propose the C2A Score, a composite measure that jointly captures correctness and constraint compliance. Results reveal a substantial gap between partial and strict satisfaction, while strong models achieve over 90% partial adherence, strict adherence remains between 39-66%. These findings highlight that trustworthy code generation requires not only correctness but also consistent adherence to developer intent.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CIFE: Code Instruction-Following Evaluation
Gunnu, Sravani
Guttula, Shanmukha
Patel, Hima
Software Engineering
Computation and Language
I.2.7; I.2.2; D.2.5
Large Language Models (LLMs) are increasingly applied to real-world code generation, where functional correctness alone is insufficient for reliable deployment, developers also expect adherence to explicit requirements for robustness, formatting, and security. Existing benchmarks primarily assess correctness through test-case execution, offering limited insight into how reliably models follow such constraints. We introduce a benchmark of 1,000 Python tasks, each paired with an average of 7 developer-specified constraints spanning 13 categories. Constraints are curated through a four-stage human-LLM pipeline to ensure they are atomic, relevant, and objective. We evaluate 14 open- and closed-source models using complementary adherence metrics and propose the C2A Score, a composite measure that jointly captures correctness and constraint compliance. Results reveal a substantial gap between partial and strict satisfaction, while strong models achieve over 90% partial adherence, strict adherence remains between 39-66%. These findings highlight that trustworthy code generation requires not only correctness but also consistent adherence to developer intent.
title CIFE: Code Instruction-Following Evaluation
topic Software Engineering
Computation and Language
I.2.7; I.2.2; D.2.5
url https://arxiv.org/abs/2512.17387