Rethinking the Evaluation of Secure Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Shih-Chieh, Xu, Jun, Tao, Guanhong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RESCUE: Retrieval Augmented Secure Code Generation
by: Shi, Jiahao, et al.
Published: (2025)
by: Shi, Jiahao, et al.
Published: (2025)
Security and Quality in LLM-Generated Code: A Multi-Language, Multi-Model Analysis
by: Kharma, Mohammed, et al.
Published: (2025)
by: Kharma, Mohammed, et al.
Published: (2025)
Evaluating and Improving the Robustness of Security Attack Detectors Generated by LLMs
by: Pasini, Samuele, et al.
Published: (2024)
by: Pasini, Samuele, et al.
Published: (2024)
Constrained Decoding for Secure Code Generation
by: Fu, Yanjun, et al.
Published: (2024)
by: Fu, Yanjun, et al.
Published: (2024)
Instruction Tuning for Secure Code Generation
by: He, Jingxuan, et al.
Published: (2024)
by: He, Jingxuan, et al.
Published: (2024)
Unintentional Security Flaws in Code: Automated Defense via Root Cause Analysis
by: Islam, Nafis Tanveer, et al.
Published: (2024)
by: Islam, Nafis Tanveer, et al.
Published: (2024)
Prompting Techniques for Secure Code Generation: A Systematic Investigation
by: Tony, Catherine, et al.
Published: (2024)
by: Tony, Catherine, et al.
Published: (2024)
HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
Deep Learning Model Security: Threats and Defenses
by: Wang, Tianyang, et al.
Published: (2024)
by: Wang, Tianyang, et al.
Published: (2024)
EditLord: Learning Code Transformation Rules for Code Editing
by: Li, Weichen, et al.
Published: (2025)
by: Li, Weichen, et al.
Published: (2025)
Enhancing Cloud Security through Topic Modelling
by: Saleh, Sabbir M., et al.
Published: (2025)
by: Saleh, Sabbir M., et al.
Published: (2025)
Expanding ML-Documentation Standards For Better Security
by: Appel, Cara Ellen
Published: (2025)
by: Appel, Cara Ellen
Published: (2025)
How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench
by: Sajadi, Amirali, et al.
Published: (2025)
by: Sajadi, Amirali, et al.
Published: (2025)
An Empirical Security Evaluation of LLM-Generated Cryptographic Rust Code
by: Elsayed, Mohamed, et al.
Published: (2026)
by: Elsayed, Mohamed, et al.
Published: (2026)
Assessing the Software Security Comprehension of Large Language Models
by: Siddiq, Mohammed Latif, et al.
Published: (2025)
by: Siddiq, Mohammed Latif, et al.
Published: (2025)
Evaluating LLM-Generated Obfuscated XSS Payloads for Machine Learning-Based Detection
by: Gabbireddy, Divyesh, et al.
Published: (2026)
by: Gabbireddy, Divyesh, et al.
Published: (2026)
"You still have to study" -- On the Security of LLM generated code
by: Goetz, Stefan, et al.
Published: (2024)
by: Goetz, Stefan, et al.
Published: (2024)
On Trojan Signatures in Large Language Models of Code
by: Hussain, Aftab, et al.
Published: (2024)
by: Hussain, Aftab, et al.
Published: (2024)
Large Language Models for Code: Security Hardening and Adversarial Testing
by: He, Jingxuan, et al.
Published: (2023)
by: He, Jingxuan, et al.
Published: (2023)
Identifier-Free Code Embedding Models for Scalable Search
by: Wolos, Eric, et al.
Published: (2026)
by: Wolos, Eric, et al.
Published: (2026)
SeBERTis: A Framework for Producing Classifiers of Security-Related Issue Reports
by: Masoumzadeh, Sogol, et al.
Published: (2025)
by: Masoumzadeh, Sogol, et al.
Published: (2025)
FINER: Enhancing State-of-the-art Classifiers with Feature Attribution to Facilitate Security Analysis
by: He, Yiling, et al.
Published: (2023)
by: He, Yiling, et al.
Published: (2023)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
LLM-based Vulnerability Discovery through the Lens of Code Metrics
by: Weissberg, Felix, et al.
Published: (2025)
by: Weissberg, Felix, et al.
Published: (2025)
Security Vulnerability Detection with Multitask Self-Instructed Fine-Tuning of Large Language Models
by: Yang, Aidan Z. H., et al.
Published: (2024)
by: Yang, Aidan Z. H., et al.
Published: (2024)
Alleviating Attack Data Scarcity: SCANIA's Experience Towards Enhancing In-Vehicle Cyber Security Measures
by: Sundfeldt, Frida, et al.
Published: (2025)
by: Sundfeldt, Frida, et al.
Published: (2025)
UniASM: Binary Code Similarity Detection without Fine-tuning
by: Gu, Yeming, et al.
Published: (2022)
by: Gu, Yeming, et al.
Published: (2022)
SecureInfer: Heterogeneous TEE-GPU Architecture for Privacy-Critical Tensors for Large Language Model Deployment
by: Nayan, Tushar, et al.
Published: (2025)
by: Nayan, Tushar, et al.
Published: (2025)
Strategic Heterogeneous Multi-Agent Architecture for Cost-Effective Code Vulnerability Detection
by: Wang, Zhaohui Geoffrey
Published: (2026)
by: Wang, Zhaohui Geoffrey
Published: (2026)
PurpCode: Reasoning for Safer Code Generation
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants
by: Štorek, Adam, et al.
Published: (2025)
by: Štorek, Adam, et al.
Published: (2025)
Code-Centric Detection of Vulnerability-Fixing Commits: A Unified Benchmark and Empirical Study
by: Loose, Nils, et al.
Published: (2026)
by: Loose, Nils, et al.
Published: (2026)
DeepCode AI Fix: Fixing Security Vulnerabilities with Large Language Models
by: Berabi, Berkay, et al.
Published: (2024)
by: Berabi, Berkay, et al.
Published: (2024)
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection
by: Charoenwet, Wachiraphan, et al.
Published: (2026)
by: Charoenwet, Wachiraphan, et al.
Published: (2026)
Software Vulnerability Prediction in Low-Resource Languages: An Empirical Study of CodeBERT and ChatGPT
by: Le, Triet H. M., et al.
Published: (2024)
by: Le, Triet H. M., et al.
Published: (2024)
AGNOMIN -- Architecture Agnostic Multi-Label Function Name Prediction
by: Achamyeleh, Yonatan Gizachew, et al.
Published: (2025)
by: Achamyeleh, Yonatan Gizachew, et al.
Published: (2025)
Automated Code-centric Software Vulnerability Assessment: How Far Are We? An Empirical Study in C/C++
by: Nguyen, Anh The, et al.
Published: (2024)
by: Nguyen, Anh The, et al.
Published: (2024)
Does Teaming-Up LLMs Improve Secure Code Generation? A Comprehensive Evaluation with Multi-LLMSecCodeEval
by: Sabir, Bushra, et al.
Published: (2026)
by: Sabir, Bushra, et al.
Published: (2026)
From Lab to Reality: A Practical Evaluation of Deep Learning Models and LLMs for Vulnerability Detection
by: Lu, Chaomeng, et al.
Published: (2025)
by: Lu, Chaomeng, et al.
Published: (2025)
REBENCH: A Procedural, Fair-by-Construction Benchmark for LLMs on Stripped-Binary Types and Names (Extended Version)
by: Won, Jun Yeon, et al.
Published: (2026)
by: Won, Jun Yeon, et al.
Published: (2026)
Similar Items
-
RESCUE: Retrieval Augmented Secure Code Generation
by: Shi, Jiahao, et al.
Published: (2025) -
Security and Quality in LLM-Generated Code: A Multi-Language, Multi-Model Analysis
by: Kharma, Mohammed, et al.
Published: (2025) -
Evaluating and Improving the Robustness of Security Attack Detectors Generated by LLMs
by: Pasini, Samuele, et al.
Published: (2024) -
Constrained Decoding for Secure Code Generation
by: Fu, Yanjun, et al.
Published: (2024) -
Instruction Tuning for Secure Code Generation
by: He, Jingxuan, et al.
Published: (2024)