Can LLM Prompting Serve as a Proxy for Static Analysis in Vulnerability Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ceka, Ira, Qiao, Feitong, Dey, Anik, Valecha, Aastha, Kaiser, Gail, Ray, Baishakhi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915490729820160
author Ceka, Ira
Qiao, Feitong
Dey, Anik
Valecha, Aastha
Kaiser, Gail
Ray, Baishakhi
author_facet Ceka, Ira
Qiao, Feitong
Dey, Anik
Valecha, Aastha
Kaiser, Gail
Ray, Baishakhi
contents Despite their remarkable success, large language models (LLMs) have shown limited ability on safety-critical code tasks such as vulnerability detection. Typically, static analysis (SA) tools, like CodeQL, CodeGuru Security, etc., are used for vulnerability detection. SA relies on predefined, manually-crafted rules for flagging various vulnerabilities. Thus, effectiveness of SA in detecting vulnerabilities depends on human experts and is known to report high error rates. In this study we investigate whether LLM prompting can be an effective alternative to these static analyzers in the partial code setting. We propose prompting strategies that integrate natural language instructions of vulnerabilities with contrastive chain-of-thought reasoning, augmented using contrastive samples from a synthetic dataset. Our findings demonstrate that security-aware prompting techniques can be effective alternatives to the laborious, hand-crafted rules of static analyzers, which often result in high false negative rates in the partial code setting. When leveraging SOTA reasoning models such as DeepSeek-R1, each of our prompting strategies exceeds the static analyzer baseline, with the best strategies improving accuracy by as much as 31.6%, F1-scores by 71.7%, pairwise accuracies by 60.4%, and reducing FNR by as much as 37.6%.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12039
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can LLM Prompting Serve as a Proxy for Static Analysis in Vulnerability Detection
Ceka, Ira
Qiao, Feitong
Dey, Anik
Valecha, Aastha
Kaiser, Gail
Ray, Baishakhi
Cryptography and Security
Artificial Intelligence
Computation and Language
Software Engineering
Despite their remarkable success, large language models (LLMs) have shown limited ability on safety-critical code tasks such as vulnerability detection. Typically, static analysis (SA) tools, like CodeQL, CodeGuru Security, etc., are used for vulnerability detection. SA relies on predefined, manually-crafted rules for flagging various vulnerabilities. Thus, effectiveness of SA in detecting vulnerabilities depends on human experts and is known to report high error rates. In this study we investigate whether LLM prompting can be an effective alternative to these static analyzers in the partial code setting. We propose prompting strategies that integrate natural language instructions of vulnerabilities with contrastive chain-of-thought reasoning, augmented using contrastive samples from a synthetic dataset. Our findings demonstrate that security-aware prompting techniques can be effective alternatives to the laborious, hand-crafted rules of static analyzers, which often result in high false negative rates in the partial code setting. When leveraging SOTA reasoning models such as DeepSeek-R1, each of our prompting strategies exceeds the static analyzer baseline, with the best strategies improving accuracy by as much as 31.6%, F1-scores by 71.7%, pairwise accuracies by 60.4%, and reducing FNR by as much as 37.6%.
title Can LLM Prompting Serve as a Proxy for Static Analysis in Vulnerability Detection
topic Cryptography and Security
Artificial Intelligence
Computation and Language
Software Engineering
url https://arxiv.org/abs/2412.12039