SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinghang, Ding, Jingzhe, Peng, Chao, Zhao, Bing, Gao, Xiang, Gao, Hongwan, Gu, Xinchen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
by: Yin, Sheng, et al.
Published: (2024)
by: Yin, Sheng, et al.
Published: (2024)
Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories
by: Yildiz, Alperen, et al.
Published: (2025)
by: Yildiz, Alperen, et al.
Published: (2025)
VULSOLVER: Vulnerability Detection via LLM-Driven Constraint Solving
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Game Rewards Vulnerabilities: Software Vulnerability Detection with Zero-Sum Game and Prototype Learning
by: Wen, Xin-Cheng, et al.
Published: (2024)
by: Wen, Xin-Cheng, et al.
Published: (2024)
Persistent Human Feedback, LLMs, and Static Analyzers for Secure Code Generation and Vulnerability Detection
by: Firouzi, Ehsan, et al.
Published: (2026)
by: Firouzi, Ehsan, et al.
Published: (2026)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
by: Shen, Chihao, et al.
Published: (2025)
by: Shen, Chihao, et al.
Published: (2025)
LogicScan: An LLM-driven Framework for Detecting Business Logic Vulnerabilities in Smart Contracts
by: Gao, Jiaqi, et al.
Published: (2026)
by: Gao, Jiaqi, et al.
Published: (2026)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
by: Chen, Qirui, et al.
Published: (2026)
by: Chen, Qirui, et al.
Published: (2026)
VulEval: Towards Repository-Level Evaluation of Software Vulnerability Detection
by: Wen, Xin-Cheng, et al.
Published: (2024)
by: Wen, Xin-Cheng, et al.
Published: (2024)
LLM-BSCVM: An LLM-Based Blockchain Smart Contract Vulnerability Management Framework
by: Jin, Yanli, et al.
Published: (2025)
by: Jin, Yanli, et al.
Published: (2025)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
by: Chen, Junkai, et al.
Published: (2025)
by: Chen, Junkai, et al.
Published: (2025)
VerilogLAVD: LLM-Aided Rule Generation for Vulnerability Detection in Verilog
by: Long, Xiang, et al.
Published: (2025)
by: Long, Xiang, et al.
Published: (2025)
LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights
by: Sheng, Ze, et al.
Published: (2025)
by: Sheng, Ze, et al.
Published: (2025)
Enhancing Code Vulnerability Detection via Vulnerability-Preserving Data Augmentation
by: Liu, Shangqing, et al.
Published: (2024)
by: Liu, Shangqing, et al.
Published: (2024)
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
by: Yang, Yixuan, et al.
Published: (2025)
by: Yang, Yixuan, et al.
Published: (2025)
GoodVibe: Security-by-Vibe for LLM-Based Code Generation
by: Thang, Maximilian, et al.
Published: (2026)
by: Thang, Maximilian, et al.
Published: (2026)
When Safe Models Merge into Danger: Exploiting Latent Vulnerabilities in LLM Fusion
by: Li, Jiaqing, et al.
Published: (2026)
by: Li, Jiaqing, et al.
Published: (2026)
Investigating Security Implications of Automatically Generated Code on the Software Supply Chain
by: Li, Xiaofan, et al.
Published: (2025)
by: Li, Xiaofan, et al.
Published: (2025)
LLM-based Vulnerable Code Augmentation: Generate or Refactor?
by: Ouchebara, Dyna Soumhane, et al.
Published: (2025)
by: Ouchebara, Dyna Soumhane, et al.
Published: (2025)
FuncVul: An Effective Function Level Vulnerability Detection Model using LLM and Code Chunk
by: Halder, Sajal, et al.
Published: (2025)
by: Halder, Sajal, et al.
Published: (2025)
Detecting Code Vulnerabilities with Heterogeneous GNN Training
by: Luo, Yu, et al.
Published: (2025)
by: Luo, Yu, et al.
Published: (2025)
Security Vulnerabilities in AI-Generated Code: A Large-Scale Analysis of Public GitHub Repositories
by: Schreiber, Maximilian, et al.
Published: (2025)
by: Schreiber, Maximilian, et al.
Published: (2025)
LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
by: Ullah, Saad, et al.
Published: (2023)
by: Ullah, Saad, et al.
Published: (2023)
HogVul: Black-box Adversarial Code Generation Framework Against LM-based Vulnerability Detectors
by: Yang, Jingxiao, et al.
Published: (2026)
by: Yang, Jingxiao, et al.
Published: (2026)
ReposVul: A Repository-Level High-Quality Vulnerability Dataset
by: Wang, Xinchen, et al.
Published: (2024)
by: Wang, Xinchen, et al.
Published: (2024)
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
by: Liu, Houjun, et al.
Published: (2026)
by: Liu, Houjun, et al.
Published: (2026)
LProtector: An LLM-driven Vulnerability Detection System
by: Sheng, Ze, et al.
Published: (2024)
by: Sheng, Ze, et al.
Published: (2024)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
by: Ying, Zonghao, et al.
Published: (2024)
by: Ying, Zonghao, et al.
Published: (2024)
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
by: Li, Rongchang, et al.
Published: (2024)
by: Li, Rongchang, et al.
Published: (2024)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
AOC-IDS: Autonomous Online Framework with Contrastive Learning for Intrusion Detection
by: Zhang, Xinchen, et al.
Published: (2024)
by: Zhang, Xinchen, et al.
Published: (2024)
VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Efficient and Universal Watermarking for LLM-Generated Code Detection
by: Li, Boquan, et al.
Published: (2024)
by: Li, Boquan, et al.
Published: (2024)
VulMCI : Code Splicing-based Pixel-row Oversampling for More Continuous Vulnerability Image Generation
by: Peng, Tao, et al.
Published: (2024)
by: Peng, Tao, et al.
Published: (2024)
Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security
by: Chua, Gabriel
Published: (2025)
by: Chua, Gabriel
Published: (2025)
VulBinLLM: LLM-powered Vulnerability Detection for Stripped Binaries
by: Hussain, Nasir, et al.
Published: (2025)
by: Hussain, Nasir, et al.
Published: (2025)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
by: Zhang, Dongsen, et al.
Published: (2025)
by: Zhang, Dongsen, et al.
Published: (2025)
CleanVul: Automatic Function-Level Vulnerability Detection in Code Commits Using LLM Heuristics
by: Li, Yikun, et al.
Published: (2024)
by: Li, Yikun, et al.
Published: (2024)
A Systematic Study of Code Obfuscation Against LLM-based Vulnerability Detection
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
CryptoGen: Secure Transformer Generation with Encrypted KV-Cache Reuse
by: Zhang, Hedong, et al.
Published: (2026)
by: Zhang, Hedong, et al.
Published: (2026)
Similar Items
-
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
by: Yin, Sheng, et al.
Published: (2024) -
Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories
by: Yildiz, Alperen, et al.
Published: (2025) -
VULSOLVER: Vulnerability Detection via LLM-Driven Constraint Solving
by: Li, Xiang, et al.
Published: (2025) -
Game Rewards Vulnerabilities: Software Vulnerability Detection with Zero-Sum Game and Prototype Learning
by: Wen, Xin-Cheng, et al.
Published: (2024) -
Persistent Human Feedback, LLMs, and Static Analyzers for Secure Code Generation and Vulnerability Detection
by: Firouzi, Ehsan, et al.
Published: (2026)