Code Roulette: How Prompt Variability Affects LLM Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paleyes, Andrei, Sendyka, Radzim, Robinson, Diana, Cabrera, Christian, Lawrence, Neil D.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914403307225088
author Paleyes, Andrei
Sendyka, Radzim
Robinson, Diana
Cabrera, Christian
Lawrence, Neil D.
author_facet Paleyes, Andrei
Sendyka, Radzim
Robinson, Diana
Cabrera, Christian
Lawrence, Neil D.
contents Code generation is one of the most active areas of application of Large Language Models (LLMs). While LLMs lower barriers to writing code and accelerate development process, the overall quality of generated programs depends on the quality of given prompts. Specifically, functionality and quality of generated code can be sensitive to user's background and familiarity with software development. It is therefore important to quantify LLM's sensitivity to variations in the input. To this end we propose an evaluation pipeline for LLM code generation with a focus on measuring sensitivity to prompt augmentations, completely agnostic to a specific programming tasks and LLMs, and thus widely applicable. We provide extensive experimental evidence illustrating utility of our method and share our code for the benefit of the community.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10204
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Code Roulette: How Prompt Variability Affects LLM Code Generation
Paleyes, Andrei
Sendyka, Radzim
Robinson, Diana
Cabrera, Christian
Lawrence, Neil D.
Software Engineering
Machine Learning
Code generation is one of the most active areas of application of Large Language Models (LLMs). While LLMs lower barriers to writing code and accelerate development process, the overall quality of generated programs depends on the quality of given prompts. Specifically, functionality and quality of generated code can be sensitive to user's background and familiarity with software development. It is therefore important to quantify LLM's sensitivity to variations in the input. To this end we propose an evaluation pipeline for LLM code generation with a focus on measuring sensitivity to prompt augmentations, completely agnostic to a specific programming tasks and LLMs, and thus widely applicable. We provide extensive experimental evidence illustrating utility of our method and share our code for the benefit of the community.
title Code Roulette: How Prompt Variability Affects LLM Code Generation
topic Software Engineering
Machine Learning
url https://arxiv.org/abs/2506.10204