How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench
Fuente:
arXiv
Saved in:
| Main Authors: | Sajadi, Amirali, Damevski, Kostadin, Chatterjee, Preetha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Personalizing Secure Programming Education with LLM-Injected Vulnerabilities
by: Frazier, Matthew, et al.
Published: (2026)
by: Frazier, Matthew, et al.
Published: (2026)
AXE: An Agentic eXploit Engine for Confirming Zero-Day Vulnerability Reports
by: Sajadi, Amirali, et al.
Published: (2026)
by: Sajadi, Amirali, et al.
Published: (2026)
Psycholinguistic Analyses in Software Engineering Text: A Systematic Literature Review
by: Sajadi, Amirali, et al.
Published: (2025)
by: Sajadi, Amirali, et al.
Published: (2025)
Do LLMs Consider Security? An Empirical Study on Responses to Programming Questions
by: Sajadi, Amirali, et al.
Published: (2025)
by: Sajadi, Amirali, et al.
Published: (2025)
Automated Repair of TEE Partitioning Issues via DSL-Guided and LLM-Assisted Patching
by: Ma, Chengyan, et al.
Published: (2026)
by: Ma, Chengyan, et al.
Published: (2026)
Accelerating Automatic Program Repair with Dual Retrieval-Augmented Fine-Tuning and Patch Generation on Large Language Models
by: Guo, Hanyang, et al.
Published: (2025)
by: Guo, Hanyang, et al.
Published: (2025)
Secure Coding with AI -- From Detection to Repair
by: Belozerov, Vladislav, et al.
Published: (2025)
by: Belozerov, Vladislav, et al.
Published: (2025)
LLM4CVE: Enabling Iterative Automated Vulnerability Repair with Large Language Models
by: Fakih, Mohamad, et al.
Published: (2025)
by: Fakih, Mohamad, et al.
Published: (2025)
A Systematic Study of LLM-Based Architectures for Automated Patching
by: Xu, Qingxiao, et al.
Published: (2026)
by: Xu, Qingxiao, et al.
Published: (2026)
Automated Repair of OpenID Connect Programs (Extended Version)
by: Rahat, Tamjid Al, et al.
Published: (2025)
by: Rahat, Tamjid Al, et al.
Published: (2025)
Toward Automated Security Risk Detection in Large Software Using Call Graph Analysis
by: Pecka, Nicholas, et al.
Published: (2025)
by: Pecka, Nicholas, et al.
Published: (2025)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
Patch-to-PoC: A Systematic Study of Agentic LLM Systems for Linux Kernel N-Day Reproduction
by: Pu, Juefei, et al.
Published: (2026)
by: Pu, Juefei, et al.
Published: (2026)
How Agentic AI Coding Assistants Become the Attacker's Shell
by: Liu, Yue, et al.
Published: (2026)
by: Liu, Yue, et al.
Published: (2026)
From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
by: Han, Junxiao, et al.
Published: (2025)
by: Han, Junxiao, et al.
Published: (2025)
Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
by: Improta, Cristina
Published: (2024)
by: Improta, Cristina
Published: (2024)
How Secure is Secure Code Generation? Adversarial Prompts Put LLM Defenses to the Test
by: Tessa, Melissa, et al.
Published: (2026)
by: Tessa, Melissa, et al.
Published: (2026)
APPATCH: Automated Adaptive Prompting Large Language Models for Real-World Software Vulnerability Patching
by: Nong, Yu, et al.
Published: (2024)
by: Nong, Yu, et al.
Published: (2024)
Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
Data Augmentation for Improving Emotion Recognition in Software Engineering Communication
by: Imran, Mia Mohammad, et al.
Published: (2022)
by: Imran, Mia Mohammad, et al.
Published: (2022)
VulKey: Automated Vulnerability Repair Guided by Domain-Specific Repair Patterns
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Root-Cause-Driven Automated Vulnerability Repair
by: Wang, Hulin, et al.
Published: (2026)
by: Wang, Hulin, et al.
Published: (2026)
How to Compare the Security of Code Written by Humans to LLM-generated Code
by: Balebako, Rebecca, et al.
Published: (2026)
by: Balebako, Rebecca, et al.
Published: (2026)
When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
by: Peng, Yibo, et al.
Published: (2025)
by: Peng, Yibo, et al.
Published: (2025)
When AI Takes the Wheel: Security Analysis of Framework-Constrained Program Generation
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
Heimdallr: Characterizing and Detecting LLM-Induced Security Risks in GitHub CI Workflows
by: Ruan, Bonan, et al.
Published: (2026)
by: Ruan, Bonan, et al.
Published: (2026)
Similar but Patched Code Considered Harmful -- The Impact of Similar but Patched Code on Recurring Vulnerability Detection and How to Remove Them
by: Tan, Zixuan, et al.
Published: (2024)
by: Tan, Zixuan, et al.
Published: (2024)
Towards Demystifying and Repairing LLM-in-the-Loop Vulnerabilities
by: Ma, Yujie, et al.
Published: (2026)
by: Ma, Yujie, et al.
Published: (2026)
StriderSPD: Structure-Guided Joint Representation Learning for Binary Security Patch Detection
by: Li, Qingyuan, et al.
Published: (2026)
by: Li, Qingyuan, et al.
Published: (2026)
AgenticCyOps: Securing Multi-Agentic AI Integration in Enterprise Cyber Operations
by: Mitra, Shaswata, et al.
Published: (2026)
by: Mitra, Shaswata, et al.
Published: (2026)
Evaluating Tool Cloning in Agentic-AI Ecosystems
by: Kim, Taein, et al.
Published: (2026)
by: Kim, Taein, et al.
Published: (2026)
LLM Security Guard for Code
by: Kavian, Arya, et al.
Published: (2024)
by: Kavian, Arya, et al.
Published: (2024)
SafeTrans: LLM-assisted Transpilation from C to Rust
by: Farrukh, Muhammad, et al.
Published: (2025)
by: Farrukh, Muhammad, et al.
Published: (2025)
Threat Modelling and Risk Analysis for Large Language Model (LLM)-Powered Applications
by: Tete, Stephen Burabari
Published: (2024)
by: Tete, Stephen Burabari
Published: (2024)
Fixing 7,400 Bugs for 1$: Cheap Crash-Site Program Repair
by: Zheng, Han, et al.
Published: (2025)
by: Zheng, Han, et al.
Published: (2025)
PatchFuzz: Patch Fuzzing for JavaScript Engines
by: Wang, Junjie, et al.
Published: (2025)
by: Wang, Junjie, et al.
Published: (2025)
VulnRepairEval: An Exploit-Based Evaluation Framework for Assessing Large Language Model Vulnerability Repair Capabilities
by: Wang, Weizhe, et al.
Published: (2025)
by: Wang, Weizhe, et al.
Published: (2025)
AIM: Automated Input Set Minimization for Metamorphic Security Testing
by: Chaleshtari, Nazanin Bayati, et al.
Published: (2024)
by: Chaleshtari, Nazanin Bayati, et al.
Published: (2024)
Empirical Study of Code Large Language Models for Binary Security Patch Detection
by: Li, Qingyuan, et al.
Published: (2025)
by: Li, Qingyuan, et al.
Published: (2025)
Similar Items
-
Towards Personalizing Secure Programming Education with LLM-Injected Vulnerabilities
by: Frazier, Matthew, et al.
Published: (2026) -
AXE: An Agentic eXploit Engine for Confirming Zero-Day Vulnerability Reports
by: Sajadi, Amirali, et al.
Published: (2026) -
Psycholinguistic Analyses in Software Engineering Text: A Systematic Literature Review
by: Sajadi, Amirali, et al.
Published: (2025) -
Do LLMs Consider Security? An Empirical Study on Responses to Programming Questions
by: Sajadi, Amirali, et al.
Published: (2025) -
Automated Repair of TEE Partitioning Issues via DSL-Guided and LLM-Assisted Patching
by: Ma, Chengyan, et al.
Published: (2026)