More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zi, Yangtian, Menon, Harshitha, Guha, Arjun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918115472834560
author Zi, Yangtian
Menon, Harshitha
Guha, Arjun
author_facet Zi, Yangtian
Menon, Harshitha
Guha, Arjun
contents State-of-the-art Large Language Models (LLMs) achieve high pass@1 on general benchmarks like HumanEval but underperform on specialized suites such as ParEval. Is this due to LLMs missing domain knowledge or insufficient prompt detail is given? To answer this, we introduce PartialOrderEval, which augments any code generation benchmark with a partial order of prompts from minimal to maximally detailed. Applying it to HumanEval and both serial and OpenMP subsets of ParEval, we measure how pass@1 scales with prompt specificity. Our experiments with Llama-3.x and Qwen2.5-Coder demonstrate varying degrees of prompt sensitivity across different tasks, and a qualitative analysis highlights explicit I/O specifications, edge-case handling, and stepwise breakdowns as the key drivers of prompt detail improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03678
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation
Zi, Yangtian
Menon, Harshitha
Guha, Arjun
Computation and Language
Machine Learning
Programming Languages
State-of-the-art Large Language Models (LLMs) achieve high pass@1 on general benchmarks like HumanEval but underperform on specialized suites such as ParEval. Is this due to LLMs missing domain knowledge or insufficient prompt detail is given? To answer this, we introduce PartialOrderEval, which augments any code generation benchmark with a partial order of prompts from minimal to maximally detailed. Applying it to HumanEval and both serial and OpenMP subsets of ParEval, we measure how pass@1 scales with prompt specificity. Our experiments with Llama-3.x and Qwen2.5-Coder demonstrate varying degrees of prompt sensitivity across different tasks, and a qualitative analysis highlights explicit I/O specifications, edge-case handling, and stepwise breakdowns as the key drivers of prompt detail improvement.
title More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation
topic Computation and Language
Machine Learning
Programming Languages
url https://arxiv.org/abs/2508.03678