SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Chihao, Dilgren, Connor, Chiniya, Purva, Griffith, Luke, Ding, Yu, Chen, Yizheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908833932115968
author Shen, Chihao
Dilgren, Connor
Chiniya, Purva
Griffith, Luke
Ding, Yu
Chen, Yizheng
author_facet Shen, Chihao
Dilgren, Connor
Chiniya, Purva
Griffith, Luke
Ding, Yu
Chen, Yizheng
contents This paper introduces SecRepoBench, a benchmark to evaluate code agents on secure code completion in real-world repositories. SecRepoBench has 318 code completion tasks in 27 C/C++ repositories, covering 15 CWEs. We evaluate 29 standalone LLMs and 15 code agents across 3 state-of-the-art agent frameworks using our benchmark. We find that state-of-the-art LLMs struggle with generating correct and secure code completions. However, code agents significantly outperform standalone LLMs. We show that SecRepoBench is more difficult than the prior state-of-the-art benchmark. Finally, our comprehensive analysis provides insights into potential directions for enhancing the ability of code agents to write correct and secure code in real-world repositories.
format Preprint
id arxiv_https___arxiv_org_abs_2504_21205
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
Shen, Chihao
Dilgren, Connor
Chiniya, Purva
Griffith, Luke
Ding, Yu
Chen, Yizheng
Cryptography and Security
Artificial Intelligence
This paper introduces SecRepoBench, a benchmark to evaluate code agents on secure code completion in real-world repositories. SecRepoBench has 318 code completion tasks in 27 C/C++ repositories, covering 15 CWEs. We evaluate 29 standalone LLMs and 15 code agents across 3 state-of-the-art agent frameworks using our benchmark. We find that state-of-the-art LLMs struggle with generating correct and secure code completions. However, code agents significantly outperform standalone LLMs. We show that SecRepoBench is more difficult than the prior state-of-the-art benchmark. Finally, our comprehensive analysis provides insights into potential directions for enhancing the ability of code agents to write correct and secure code in real-world repositories.
title SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2504.21205