Why LLMs Fail: A Failure Analysis and Partial Success Measurement for Automated Security Patch Generation
Fuente:
arXiv
Saved in:
| Main Author: | Al-Maamari, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can You Trust Your Copilot? A Privacy Scorecard for AI Coding Assistants
by: AL-Maamari, Amir
Published: (2025)
by: AL-Maamari, Amir
Published: (2025)
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
by: Bhatt, Manish, et al.
Published: (2026)
by: Bhatt, Manish, et al.
Published: (2026)
On The Dangers of Poisoned LLMs In Security Automation
by: Karlsen, Patrick, et al.
Published: (2025)
by: Karlsen, Patrick, et al.
Published: (2025)
On Automating Security Policies with Contemporary LLMs
by: Saura, Pablo Fernández, et al.
Published: (2025)
by: Saura, Pablo Fernández, et al.
Published: (2025)
STAF: Leveraging LLMs for Automated Attack Tree-Based Security Test Generation
by: Khule, Tanmay, et al.
Published: (2025)
by: Khule, Tanmay, et al.
Published: (2025)
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
by: Abdelnabi, Sahar, et al.
Published: (2026)
by: Abdelnabi, Sahar, et al.
Published: (2026)
SecureFalcon: Are We There Yet in Automated Software Vulnerability Detection with LLMs?
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
Just-in-Time Detection of Silent Security Patches
by: Tang, Xunzhu, et al.
Published: (2023)
by: Tang, Xunzhu, et al.
Published: (2023)
Securing LLM-Generated Embedded Firmware through AI Agent-Driven Validation and Patching
by: Abtahi, Seyed Moein, et al.
Published: (2025)
by: Abtahi, Seyed Moein, et al.
Published: (2025)
Retrieval-Augmented LLMs for Security Incident Analysis
by: Cadet, Xavier, et al.
Published: (2026)
by: Cadet, Xavier, et al.
Published: (2026)
Cracking IoT Security: Can LLMs Outsmart Static Analysis Tools?
by: Quantrill, Jason, et al.
Published: (2026)
by: Quantrill, Jason, et al.
Published: (2026)
Ocassionally Secure: A Comparative Analysis of Code Generation Assistants
by: Elgedawy, Ran, et al.
Published: (2024)
by: Elgedawy, Ran, et al.
Published: (2024)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
by: Chen, Qirui, et al.
Published: (2026)
by: Chen, Qirui, et al.
Published: (2026)
Enhancing Source Code Security with LLMs: Demystifying The Challenges and Generating Reliable Repairs
by: Islam, Nafis Tanveer, et al.
Published: (2024)
by: Islam, Nafis Tanveer, et al.
Published: (2024)
Towards Compositional Generalization in LLMs for Smart Contract Security: A Case Study on Reentrancy Vulnerabilities
by: Zhou, Ying, et al.
Published: (2026)
by: Zhou, Ying, et al.
Published: (2026)
Automated Consistency Analysis of LLMs
by: Patwardhan, Aditya, et al.
Published: (2025)
by: Patwardhan, Aditya, et al.
Published: (2025)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
by: Shang, Zhengchun, et al.
Published: (2025)
by: Shang, Zhengchun, et al.
Published: (2025)
DUALGUAGE: Automated Joint Security-Functionality Benchmarking for Secure Code Generation
by: Pathak, Abhijeet, et al.
Published: (2025)
by: Pathak, Abhijeet, et al.
Published: (2025)
LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights
by: Sheng, Ze, et al.
Published: (2025)
by: Sheng, Ze, et al.
Published: (2025)
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
by: Zloczower, Itay, et al.
Published: (2026)
by: Zloczower, Itay, et al.
Published: (2026)
Machine Learning-Based Security Policy Analysis
by: Jain, Krish, et al.
Published: (2024)
by: Jain, Krish, et al.
Published: (2024)
AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
From Theory to Practice: Code Generation Using LLMs for CAPEC and CWE Frameworks
by: Shahzad, Murtuza, et al.
Published: (2026)
by: Shahzad, Murtuza, et al.
Published: (2026)
When Efficiency Backfires: Cascading LLMs Trigger Cascade Failure under Adversarial Attack
by: Sun, Zehan, et al.
Published: (2026)
by: Sun, Zehan, et al.
Published: (2026)
Measuring Safety Alignment Effects in Autonomous Security Agents
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Deep Learning Under Siege: Identifying Security Vulnerabilities and Risk Mitigation Strategies
by: Al-Karaki, Jamal, et al.
Published: (2024)
by: Al-Karaki, Jamal, et al.
Published: (2024)
PenTest++: Elevating Ethical Hacking with AI and Automation
by: Al-Sinani, Haitham S., et al.
Published: (2025)
by: Al-Sinani, Haitham S., et al.
Published: (2025)
PatUntrack: Automated Generating Patch Examples for Issue Reports without Tracked Insecure Code
by: Jiang, Ziyou, et al.
Published: (2024)
by: Jiang, Ziyou, et al.
Published: (2024)
Contextualized AI for Cyber Defense: An Automated Survey using LLMs
by: Haryanto, Christoforus Yoga, et al.
Published: (2024)
by: Haryanto, Christoforus Yoga, et al.
Published: (2024)
Distillability of LLM Security Logic: Predicting Attack Success Rate of Outline Filling Attack via Ranking Regression
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
WebTrap Park: An Automated Platform for Systematic Security Evaluation of Web Agents
by: Wu, Xinyi, et al.
Published: (2026)
by: Wu, Xinyi, et al.
Published: (2026)
A Systematic Evaluation of Parameter-Efficient Fine-Tuning Methods for the Security of Code LLMs
by: Lee, Kiho, et al.
Published: (2025)
by: Lee, Kiho, et al.
Published: (2025)
Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
FedSecurity: Benchmarking Attacks and Defenses in Federated Learning and Federated LLMs
by: Han, Shanshan, et al.
Published: (2023)
by: Han, Shanshan, et al.
Published: (2023)
Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis
by: Li, Zhiyuan, et al.
Published: (2026)
by: Li, Zhiyuan, et al.
Published: (2026)
Quantifying Security Vulnerabilities: A Metric-Driven Security Analysis of Gaps in Current AI Standards
by: Madhavan, Keerthana, et al.
Published: (2025)
by: Madhavan, Keerthana, et al.
Published: (2025)
Security of and by Generative AI platforms
by: Hayagreevan, Hari, et al.
Published: (2024)
by: Hayagreevan, Hari, et al.
Published: (2024)
Secure Multiparty Generative AI
by: Shrestha, Manil, et al.
Published: (2024)
by: Shrestha, Manil, et al.
Published: (2024)
Automating Security Audit Using Large Language Model based Agent: An Exploration Experiment
by: Chin, Jia Hui, et al.
Published: (2025)
by: Chin, Jia Hui, et al.
Published: (2025)
Similar Items
-
Can You Trust Your Copilot? A Privacy Scorecard for AI Coding Assistants
by: AL-Maamari, Amir
Published: (2025) -
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
by: Bhatt, Manish, et al.
Published: (2026) -
On The Dangers of Poisoned LLMs In Security Automation
by: Karlsen, Patrick, et al.
Published: (2025) -
On Automating Security Policies with Contemporary LLMs
by: Saura, Pablo Fernández, et al.
Published: (2025) -
STAF: Leveraging LLMs for Automated Attack Tree-Based Security Test Generation
by: Khule, Tanmay, et al.
Published: (2025)