Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.24171 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916041036136448 |
|---|---|
| author | Camarato, Steffen J. Hmaiti, Yahya Ghadamian, Mandana Mohaisen, David |
| author_facet | Camarato, Steffen J. Hmaiti, Yahya Ghadamian, Mandana Mohaisen, David |
| contents | Large language models are increasingly used for vulnerability detection, yet their reliability under different prompt formulations remains uncharacterized. We present PromptAudit, a controlled evaluation framework that isolates prompt effects by fixing the dataset, decoding, and parsing while varying only the prompting strategy. Using five prompting strategies across five open-weight models on 1,000 CVEs (6,074 code samples spanning 16 programming languages), we evaluate accuracy, recall, abstention, coverage, and effective F1. We find that standard chain-of-thought prompting achieves the strongest overall operational performance, while few-shot prompting provides model-dependent benefits that are most pronounced for prompt-sensitive models. In contrast, adaptive chain-of-thought frequently suppresses recall and self-consistency induces excessive abstention, sharply reducing effective performance. These results show that vulnerability detection behavior is jointly determined by the model and the prompt, and that prompt sensitivity is a first-class system property that must be explicitly characterized in evaluation and deployment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_24171 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection Camarato, Steffen J. Hmaiti, Yahya Ghadamian, Mandana Mohaisen, David Machine Learning Artificial Intelligence Large language models are increasingly used for vulnerability detection, yet their reliability under different prompt formulations remains uncharacterized. We present PromptAudit, a controlled evaluation framework that isolates prompt effects by fixing the dataset, decoding, and parsing while varying only the prompting strategy. Using five prompting strategies across five open-weight models on 1,000 CVEs (6,074 code samples spanning 16 programming languages), we evaluate accuracy, recall, abstention, coverage, and effective F1. We find that standard chain-of-thought prompting achieves the strongest overall operational performance, while few-shot prompting provides model-dependent benefits that are most pronounced for prompt-sensitive models. In contrast, adaptive chain-of-thought frequently suppresses recall and self-consistency induces excessive abstention, sharply reducing effective performance. These results show that vulnerability detection behavior is jointly determined by the model and the prompt, and that prompt sensitivity is a first-class system property that must be explicitly characterized in evaluation and deployment. |
| title | PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2605.24171 |