RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yanlin, Zhang, Ziyao, Wang, Chong, Xu, Xinyi, Liu, Mingwei, Wang, Yong, Chen, Jiachi, Zheng, Zibin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915763458146304
author Wang, Yanlin
Zhang, Ziyao
Wang, Chong
Xu, Xinyi
Liu, Mingwei
Wang, Yong
Chen, Jiachi
Zheng, Zibin
author_facet Wang, Yanlin
Zhang, Ziyao
Wang, Chong
Xu, Xinyi
Liu, Mingwei
Wang, Yong
Chen, Jiachi
Zheng, Zibin
contents Large Language Models (LLMs) have demonstrated remarkable capabilities in code generation, but their proficiency in producing secure code remains a critical, under-explored area. Existing benchmarks often fall short by relying on synthetic vulnerabilities or evaluating functional correctness in isolation, failing to capture the complex interplay between functionality and security found in real-world software. To address this gap, we introduce RealSec-bench, a new benchmark for secure code generation meticulously constructed from real-world, high-risk Java repositories. Our methodology employs a multi-stage pipeline that combines systematic SAST scanning with CodeQL, LLM-based false positive elimination, and rigorous human expert validation. The resulting benchmark contains 105 instances grounded in real-word repository contexts, spanning 19 Common Weakness Enumeration (CWE) types and exhibiting a wide diversity of data flow complexities, including vulnerabilities with up to 34-hop inter-procedural dependencies. Using RealSec-bench, we conduct an extensive empirical study on 5 popular LLMs. We introduce a novel composite metric, SecurePass@K, to assess both functional correctness and security simultaneously. We find that while Retrieval-Augmented Generation (RAG) techniques can improve functional correctness, they provide negligible benefits to security. Furthermore, explicitly prompting models with general security guidelines often leads to compilation failures, harming functional correctness without reliably preventing vulnerabilities. Our work highlights the gap between functional and secure code generation in current LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22706
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
Wang, Yanlin
Zhang, Ziyao
Wang, Chong
Xu, Xinyi
Liu, Mingwei
Wang, Yong
Chen, Jiachi
Zheng, Zibin
Cryptography and Security
Software Engineering
Large Language Models (LLMs) have demonstrated remarkable capabilities in code generation, but their proficiency in producing secure code remains a critical, under-explored area. Existing benchmarks often fall short by relying on synthetic vulnerabilities or evaluating functional correctness in isolation, failing to capture the complex interplay between functionality and security found in real-world software. To address this gap, we introduce RealSec-bench, a new benchmark for secure code generation meticulously constructed from real-world, high-risk Java repositories. Our methodology employs a multi-stage pipeline that combines systematic SAST scanning with CodeQL, LLM-based false positive elimination, and rigorous human expert validation. The resulting benchmark contains 105 instances grounded in real-word repository contexts, spanning 19 Common Weakness Enumeration (CWE) types and exhibiting a wide diversity of data flow complexities, including vulnerabilities with up to 34-hop inter-procedural dependencies. Using RealSec-bench, we conduct an extensive empirical study on 5 popular LLMs. We introduce a novel composite metric, SecurePass@K, to assess both functional correctness and security simultaneously. We find that while Retrieval-Augmented Generation (RAG) techniques can improve functional correctness, they provide negligible benefits to security. Furthermore, explicitly prompting models with general security guidelines often leads to compilation failures, harming functional correctness without reliably preventing vulnerabilities. Our work highlights the gap between functional and secure code generation in current LLMs.
title RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
topic Cryptography and Security
Software Engineering
url https://arxiv.org/abs/2601.22706