SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908833932115968 |
|---|---|
| author | Shen, Chihao Dilgren, Connor Chiniya, Purva Griffith, Luke Ding, Yu Chen, Yizheng |
| author_facet | Shen, Chihao Dilgren, Connor Chiniya, Purva Griffith, Luke Ding, Yu Chen, Yizheng |
| contents | This paper introduces SecRepoBench, a benchmark to evaluate code agents on secure code completion in real-world repositories. SecRepoBench has 318 code completion tasks in 27 C/C++ repositories, covering 15 CWEs. We evaluate 29 standalone LLMs and 15 code agents across 3 state-of-the-art agent frameworks using our benchmark. We find that state-of-the-art LLMs struggle with generating correct and secure code completions. However, code agents significantly outperform standalone LLMs. We show that SecRepoBench is more difficult than the prior state-of-the-art benchmark. Finally, our comprehensive analysis provides insights into potential directions for enhancing the ability of code agents to write correct and secure code in real-world repositories. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_21205 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories Shen, Chihao Dilgren, Connor Chiniya, Purva Griffith, Luke Ding, Yu Chen, Yizheng Cryptography and Security Artificial Intelligence This paper introduces SecRepoBench, a benchmark to evaluate code agents on secure code completion in real-world repositories. SecRepoBench has 318 code completion tasks in 27 C/C++ repositories, covering 15 CWEs. We evaluate 29 standalone LLMs and 15 code agents across 3 state-of-the-art agent frameworks using our benchmark. We find that state-of-the-art LLMs struggle with generating correct and secure code completions. However, code agents significantly outperform standalone LLMs. We show that SecRepoBench is more difficult than the prior state-of-the-art benchmark. Finally, our comprehensive analysis provides insights into potential directions for enhancing the ability of code agents to write correct and secure code in real-world repositories. |
| title | SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2504.21205 |