SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Chihao, Dilgren, Connor, Chiniya, Purva, Griffith, Luke, Ding, Yu, Chen, Yizheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
by: Chen, Qirui, et al.
Published: (2026)
by: Chen, Qirui, et al.
Published: (2026)
SecCodePRM: A Process Reward Model for Code Security
by: Yu, Weichen, et al.
Published: (2026)
by: Yu, Weichen, et al.
Published: (2026)
Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories
by: Yildiz, Alperen, et al.
Published: (2025)
by: Yildiz, Alperen, et al.
Published: (2025)
Constrained Decoding for Secure Code Generation
by: Fu, Yanjun, et al.
Published: (2024)
by: Fu, Yanjun, et al.
Published: (2024)
SecCodeBench-V2 Technical Report
by: Chen, Longfei, et al.
Published: (2026)
by: Chen, Longfei, et al.
Published: (2026)
SecReEvalBench: A Multi-turned Security Resilience Evaluation Benchmark for Large Language Models
by: Cui, Huining, et al.
Published: (2025)
by: Cui, Huining, et al.
Published: (2025)
SecPI: Secure Code Generation with Reasoning Models via Security Reasoning Internalization
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
CIBER: A Comprehensive Benchmark for Security Evaluation of Code Interpreter Agents
by: Ba, Lei, et al.
Published: (2026)
by: Ba, Lei, et al.
Published: (2026)
ProSec: Fortifying Code LLMs with Proactive Security Alignment
by: Xu, Xiangzhe, et al.
Published: (2024)
by: Xu, Xiangzhe, et al.
Published: (2024)
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
by: Li, Xinghang, et al.
Published: (2025)
by: Li, Xinghang, et al.
Published: (2025)
CellularSpecSec-Bench: A Staged Benchmark for Evidence-Grounded Interpretation and Security Reasoning over 3GPP Specifications
by: Xie, Ke, et al.
Published: (2026)
by: Xie, Ke, et al.
Published: (2026)
SecRef*: Securely Sharing Mutable References Between Verified and Unverified Code in F*
by: Andrici, Cezar-Constantin, et al.
Published: (2025)
by: Andrici, Cezar-Constantin, et al.
Published: (2025)
ReposVul: A Repository-Level High-Quality Vulnerability Dataset
by: Wang, Xinchen, et al.
Published: (2024)
by: Wang, Xinchen, et al.
Published: (2024)
Security Vulnerabilities in AI-Generated Code: A Large-Scale Analysis of Public GitHub Repositories
by: Schreiber, Maximilian, et al.
Published: (2025)
by: Schreiber, Maximilian, et al.
Published: (2025)
SecGoal: A Benchmark for Extracting Formalizable Security Goals from Protocol Documents
by: Huang, Dawei, et al.
Published: (2026)
by: Huang, Dawei, et al.
Published: (2026)
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
by: Jing, Pengfei, et al.
Published: (2024)
by: Jing, Pengfei, et al.
Published: (2024)
Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
by: Luo, Weidi, et al.
Published: (2025)
by: Luo, Weidi, et al.
Published: (2025)
Security Attacks on LLM-based Code Completion Tools
by: Cheng, Wen, et al.
Published: (2024)
by: Cheng, Wen, et al.
Published: (2024)
Detecting Hard-Coded Credentials in Software Repositories via LLMs
by: Biringa, Chidera, et al.
Published: (2025)
by: Biringa, Chidera, et al.
Published: (2025)
SecIC3: Customizing IC3 for Hardware Security Verification
by: Tan, Qinhan, et al.
Published: (2026)
by: Tan, Qinhan, et al.
Published: (2026)
Locus: Agentic Predicate Synthesis for Directed Fuzzing
by: Zhu, Jie, et al.
Published: (2025)
by: Zhu, Jie, et al.
Published: (2025)
RepoMark: A Data-Usage Auditing Framework for Code Large Language Models
by: Qu, Wenjie, et al.
Published: (2025)
by: Qu, Wenjie, et al.
Published: (2025)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
by: Chen, Junkai, et al.
Published: (2025)
by: Chen, Junkai, et al.
Published: (2025)
Hooked: A Real-World Study on QR Code Phishing
by: Geisler, Marvin, et al.
Published: (2024)
by: Geisler, Marvin, et al.
Published: (2024)
SecDTD: Dynamic Token Drop for Secure Transformers Inference
by: Cai, Yifei, et al.
Published: (2026)
by: Cai, Yifei, et al.
Published: (2026)
CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models
by: Wahréus, Johan, et al.
Published: (2025)
by: Wahréus, Johan, et al.
Published: (2025)
Generalized Adversarial Code-Suggestions: Exploiting Contexts of LLM-based Code-Completion
by: Rubel, Karl, et al.
Published: (2024)
by: Rubel, Karl, et al.
Published: (2024)
AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
SecureCodeRL: Security-Aware Reinforcement Learning for Code Generation with Partial-Credit Rewards
by: Sijwali, Suryansh Singh, et al.
Published: (2026)
by: Sijwali, Suryansh Singh, et al.
Published: (2026)
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
by: Lee, Hwiwon, et al.
Published: (2025)
by: Lee, Hwiwon, et al.
Published: (2025)
Your Code Secret Belongs to Me: Neural Code Completion Tools Can Memorize Hard-Coded Credentials
by: Huang, Yizhan, et al.
Published: (2023)
by: Huang, Yizhan, et al.
Published: (2023)
CellSecInspector: Safeguarding Cellular Networks via Automated Security Analysis on Specifications
by: Xie, Ke, et al.
Published: (2025)
by: Xie, Ke, et al.
Published: (2025)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
by: Yuan, He Yang, et al.
Published: (2026)
by: Yuan, He Yang, et al.
Published: (2026)
Using Real-world Bug Bounty Programs in Secure Coding Course: Experience Report
by: Malinka, Kamil, et al.
Published: (2024)
by: Malinka, Kamil, et al.
Published: (2024)
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
by: Saha, Shoumik, et al.
Published: (2025)
by: Saha, Shoumik, et al.
Published: (2025)
SecFSM: Knowledge Graph-Guided Verilog Code Generation for Secure Finite State Machines in Systems-on-Chip
by: Hu, Ziteng, et al.
Published: (2025)
by: Hu, Ziteng, et al.
Published: (2025)
RealVuln: Benchmarking Rule-Based, General-Purpose LLM, and Security-Specialized Scanners on Real-World Code
by: Pellew, John, et al.
Published: (2026)
by: Pellew, John, et al.
Published: (2026)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
Similar Items
-
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026) -
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
by: Chen, Qirui, et al.
Published: (2026) -
SecCodePRM: A Process Reward Model for Code Security
by: Yu, Weichen, et al.
Published: (2026) -
Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories
by: Yildiz, Alperen, et al.
Published: (2025) -
Constrained Decoding for Secure Code Generation
by: Fu, Yanjun, et al.
Published: (2024)