Saved in:
Bibliographic Details
Main Authors: Camarato, Steffen J., Hmaiti, Yahya, Ghadamian, Mandana, Mohaisen, David
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.24171
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916041036136448
author Camarato, Steffen J.
Hmaiti, Yahya
Ghadamian, Mandana
Mohaisen, David
author_facet Camarato, Steffen J.
Hmaiti, Yahya
Ghadamian, Mandana
Mohaisen, David
contents Large language models are increasingly used for vulnerability detection, yet their reliability under different prompt formulations remains uncharacterized. We present PromptAudit, a controlled evaluation framework that isolates prompt effects by fixing the dataset, decoding, and parsing while varying only the prompting strategy. Using five prompting strategies across five open-weight models on 1,000 CVEs (6,074 code samples spanning 16 programming languages), we evaluate accuracy, recall, abstention, coverage, and effective F1. We find that standard chain-of-thought prompting achieves the strongest overall operational performance, while few-shot prompting provides model-dependent benefits that are most pronounced for prompt-sensitive models. In contrast, adaptive chain-of-thought frequently suppresses recall and self-consistency induces excessive abstention, sharply reducing effective performance. These results show that vulnerability detection behavior is jointly determined by the model and the prompt, and that prompt sensitivity is a first-class system property that must be explicitly characterized in evaluation and deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24171
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection
Camarato, Steffen J.
Hmaiti, Yahya
Ghadamian, Mandana
Mohaisen, David
Machine Learning
Artificial Intelligence
Large language models are increasingly used for vulnerability detection, yet their reliability under different prompt formulations remains uncharacterized. We present PromptAudit, a controlled evaluation framework that isolates prompt effects by fixing the dataset, decoding, and parsing while varying only the prompting strategy. Using five prompting strategies across five open-weight models on 1,000 CVEs (6,074 code samples spanning 16 programming languages), we evaluate accuracy, recall, abstention, coverage, and effective F1. We find that standard chain-of-thought prompting achieves the strongest overall operational performance, while few-shot prompting provides model-dependent benefits that are most pronounced for prompt-sensitive models. In contrast, adaptive chain-of-thought frequently suppresses recall and self-consistency induces excessive abstention, sharply reducing effective performance. These results show that vulnerability detection behavior is jointly determined by the model and the prompt, and that prompt sensitivity is a first-class system property that must be explicitly characterized in evaluation and deployment.
title PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.24171