SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
Fuente:
arXiv
Saved in:
| Main Author: | Thornton, Scott |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
by: Wahed, Muntasir, et al.
Published: (2025)
by: Wahed, Muntasir, et al.
Published: (2025)
HexaCoder: Secure Code Generation via Oracle-Guided Synthetic Training Data
by: Hajipour, Hossein, et al.
Published: (2024)
by: Hajipour, Hossein, et al.
Published: (2024)
Security Degradation in Iterative AI Code Generation -- A Systematic Analysis of the Paradox
by: Shukla, Shivani, et al.
Published: (2025)
by: Shukla, Shivani, et al.
Published: (2025)
AutoBaxBuilder: Bootstrapping Code Security Benchmarking
by: von Arx, Tobias, et al.
Published: (2025)
by: von Arx, Tobias, et al.
Published: (2025)
Automated Software Vulnerability Static Code Analysis Using Generative Pre-Trained Transformer Models
by: Pelofske, Elijah, et al.
Published: (2024)
by: Pelofske, Elijah, et al.
Published: (2024)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
by: Tong, Terry, et al.
Published: (2024)
by: Tong, Terry, et al.
Published: (2024)
A Fast, Reliable, and Secure Programming Language for LLM Agents with Code Actions
by: Mell, Stephen, et al.
Published: (2025)
by: Mell, Stephen, et al.
Published: (2025)
Enhancing Reliability in LLM-Based Secure Code Generation
by: Kharma, Mohammed F., et al.
Published: (2026)
by: Kharma, Mohammed F., et al.
Published: (2026)
Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model
by: Wu, Tianyi, et al.
Published: (2026)
by: Wu, Tianyi, et al.
Published: (2026)
Gandalf the Red: Adaptive Security for LLMs
by: Pfister, Niklas, et al.
Published: (2025)
by: Pfister, Niklas, et al.
Published: (2025)
Generative AI Security: Challenges and Countermeasures
by: Zhu, Banghua, et al.
Published: (2024)
by: Zhu, Banghua, et al.
Published: (2024)
SecureBreak -- A dataset towards safe and secure models
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
by: Reddy, Aashray, et al.
Published: (2025)
by: Reddy, Aashray, et al.
Published: (2025)
An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods
by: Kharma, Mohammed, et al.
Published: (2026)
by: Kharma, Mohammed, et al.
Published: (2026)
Constrained Decoding for Secure Code Generation
by: Fu, Yanjun, et al.
Published: (2024)
by: Fu, Yanjun, et al.
Published: (2024)
Instruction Tuning for Secure Code Generation
by: He, Jingxuan, et al.
Published: (2024)
by: He, Jingxuan, et al.
Published: (2024)
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
by: Wang, Guanchu, et al.
Published: (2024)
by: Wang, Guanchu, et al.
Published: (2024)
SecEncoder: Logs are All You Need in Security
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
SentinelLMs: Encrypted Input Adaptation and Fine-tuning of Language Models for Private and Secure Inference
by: Mishra, Abhijit, et al.
Published: (2023)
by: Mishra, Abhijit, et al.
Published: (2023)
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
by: Ren, Qibing, et al.
Published: (2024)
by: Ren, Qibing, et al.
Published: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis
by: Mostafa, Ahmed, et al.
Published: (2025)
by: Mostafa, Ahmed, et al.
Published: (2025)
Semantic Chameleon: Corpus-Dependent Poisoning Attacks and Defenses in RAG Systems
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
From Vulnerabilities to Remediation: A Systematic Literature Review of LLMs in Code Security
by: Basic, Enna, et al.
Published: (2024)
by: Basic, Enna, et al.
Published: (2024)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
by: Guo, Chuan, et al.
Published: (2026)
by: Guo, Chuan, et al.
Published: (2026)
Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG
by: Singh, Inderjeet, et al.
Published: (2026)
by: Singh, Inderjeet, et al.
Published: (2026)
Noise Contrastive Estimation-based Matching Framework for Low-Resource Security Attack Pattern Recognition
by: Nguyen, Tu, et al.
Published: (2024)
by: Nguyen, Tu, et al.
Published: (2024)
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
by: Li, Xueyi, et al.
Published: (2026)
by: Li, Xueyi, et al.
Published: (2026)
Prompting Techniques for Secure Code Generation: A Systematic Investigation
by: Tony, Catherine, et al.
Published: (2024)
by: Tony, Catherine, et al.
Published: (2024)
Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign
by: Zhang, Ruisi, et al.
Published: (2025)
by: Zhang, Ruisi, et al.
Published: (2025)
BaxBench: Can LLMs Generate Correct and Secure Backends?
by: Vero, Mark, et al.
Published: (2025)
by: Vero, Mark, et al.
Published: (2025)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
by: Salahuddin, Salahuddin, et al.
Published: (2025)
by: Salahuddin, Salahuddin, et al.
Published: (2025)
Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
by: Dang, Trung Cuong, et al.
Published: (2025)
by: Dang, Trung Cuong, et al.
Published: (2025)
A Security Risk Taxonomy for Prompt-Based Interaction With Large Language Models
by: Derner, Erik, et al.
Published: (2023)
by: Derner, Erik, et al.
Published: (2023)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
by: Banayeeanzade, Amin, et al.
Published: (2026)
by: Banayeeanzade, Amin, et al.
Published: (2026)
Survival of the Safest: Towards Secure Prompt Optimization through Interleaved Multi-Objective Evolution
by: Sinha, Ankita, et al.
Published: (2024)
by: Sinha, Ankita, et al.
Published: (2024)
TableGuard -- Securing Structured & Unstructured Data
by: Sharma, Anantha, et al.
Published: (2024)
by: Sharma, Anantha, et al.
Published: (2024)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
by: Rahman, Salman, et al.
Published: (2025)
by: Rahman, Salman, et al.
Published: (2025)
Similar Items
-
Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis
by: Thornton, Scott
Published: (2026) -
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
by: Wahed, Muntasir, et al.
Published: (2025) -
HexaCoder: Secure Code Generation via Oracle-Guided Synthetic Training Data
by: Hajipour, Hossein, et al.
Published: (2024) -
Security Degradation in Iterative AI Code Generation -- A Systematic Analysis of the Paradox
by: Shukla, Shivani, et al.
Published: (2025) -
AutoBaxBuilder: Bootstrapping Code Security Benchmarking
by: von Arx, Tobias, et al.
Published: (2025)