Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Xueying, Zheng, Geng, Wang, Kaixin, Zou, Yi, Wang, Yujia, Deng, Wentai, Feng, Jiayi, Liu, Mingwei, Chen, Bihuan, Peng, Xin, Ma, Tao, Lou, Yiling
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909650761285632
author Du, Xueying
Zheng, Geng
Wang, Kaixin
Zou, Yi
Wang, Yujia
Deng, Wentai
Feng, Jiayi
Liu, Mingwei
Chen, Bihuan
Peng, Xin
Ma, Tao
Lou, Yiling
author_facet Du, Xueying
Zheng, Geng
Wang, Kaixin
Zou, Yi
Wang, Yujia
Deng, Wentai
Feng, Jiayi
Liu, Mingwei
Chen, Bihuan
Peng, Xin
Ma, Tao
Lou, Yiling
contents Although LLMs have shown promising potential in vulnerability detection, this study reveals their limitations in distinguishing between vulnerable and similar-but-benign patched code (only 0.06 - 0.14 accuracy). It shows that LLMs struggle to capture the root causes of vulnerabilities during vulnerability detection. To address this challenge, we propose enhancing LLMs with multi-dimensional vulnerability knowledge distilled from historical vulnerabilities and fixes. We design a novel knowledge-level Retrieval-Augmented Generation framework Vul-RAG, which improves LLMs with an accuracy increase of 16% - 24% in identifying vulnerable and patched code. Additionally, vulnerability knowledge generated by Vul-RAG can further (1) serve as high-quality explanations to improve manual detection accuracy (from 60% to 77%), and (2) detect 10 previously-unknown bugs in the recent Linux kernel release with 6 assigned CVEs.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11147
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG
Du, Xueying
Zheng, Geng
Wang, Kaixin
Zou, Yi
Wang, Yujia
Deng, Wentai
Feng, Jiayi
Liu, Mingwei
Chen, Bihuan
Peng, Xin
Ma, Tao
Lou, Yiling
Software Engineering
Artificial Intelligence
Although LLMs have shown promising potential in vulnerability detection, this study reveals their limitations in distinguishing between vulnerable and similar-but-benign patched code (only 0.06 - 0.14 accuracy). It shows that LLMs struggle to capture the root causes of vulnerabilities during vulnerability detection. To address this challenge, we propose enhancing LLMs with multi-dimensional vulnerability knowledge distilled from historical vulnerabilities and fixes. We design a novel knowledge-level Retrieval-Augmented Generation framework Vul-RAG, which improves LLMs with an accuracy increase of 16% - 24% in identifying vulnerable and patched code. Additionally, vulnerability knowledge generated by Vul-RAG can further (1) serve as high-quality explanations to improve manual detection accuracy (from 60% to 77%), and (2) detect 10 previously-unknown bugs in the recent Linux kernel release with 6 assigned CVEs.
title Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2406.11147