Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Saha, Shoumik, Chen, Jifan, Mayers, Sam, Gouda, Sanjay Krishna, Wang, Zijian, Kumar, Varun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
by: Saha, Shoumik, et al.
Published: (2026)
by: Saha, Shoumik, et al.
Published: (2026)
SoK: Understanding (New) Security Issues Across AI4Code Use Cases
by: Wu, Qilong, et al.
Published: (2025)
by: Wu, Qilong, et al.
Published: (2025)
Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
by: Xie, Yuchong, et al.
Published: (2025)
by: Xie, Yuchong, et al.
Published: (2025)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
by: Shen, Chihao, et al.
Published: (2025)
by: Shen, Chihao, et al.
Published: (2025)
ESAA-Security: An Event-Sourced, Verifiable Architecture for Agent-Assisted Security Audits of AI-Generated Code
by: Filho, Elzo Brito dos Santos
Published: (2026)
by: Filho, Elzo Brito dos Santos
Published: (2026)
Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
by: Kim, Jungin, et al.
Published: (2025)
by: Kim, Jungin, et al.
Published: (2025)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
by: Shang, Zhengchun, et al.
Published: (2025)
by: Shang, Zhengchun, et al.
Published: (2025)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
by: Xiang, Chong, et al.
Published: (2026)
by: Xiang, Chong, et al.
Published: (2026)
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
Securing AI Agents Against Prompt Injection Attacks
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
On Jailbreaking Quantized Language Models Through Fault Injection Attacks
by: Zahran, Noureldin, et al.
Published: (2025)
by: Zahran, Noureldin, et al.
Published: (2025)
EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion
by: Liang, Zhen, et al.
Published: (2025)
by: Liang, Zhen, et al.
Published: (2025)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
by: Chen, Qirui, et al.
Published: (2026)
by: Chen, Qirui, et al.
Published: (2026)
When Developer Aid Becomes Security Debt: A Systematic Analysis of Insecure Behaviors in LLM Coding Agents
by: Kozak, Matous, et al.
Published: (2025)
by: Kozak, Matous, et al.
Published: (2025)
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
by: Wang, Xiangwen, et al.
Published: (2026)
by: Wang, Xiangwen, et al.
Published: (2026)
AttackQA: Development and Adoption of a Dataset for Assisting Cybersecurity Operations using Fine-tuned and Open-Source LLMs
by: Krishna, Varun Badrinath
Published: (2024)
by: Krishna, Varun Badrinath
Published: (2024)
Fast Adversarial Attacks on Language Models In One GPU Minute
by: Sadasivan, Vinu Sankar, et al.
Published: (2024)
by: Sadasivan, Vinu Sankar, et al.
Published: (2024)
SecureCodeRL: Security-Aware Reinforcement Learning for Code Generation with Partial-Credit Rewards
by: Sijwali, Suryansh Singh, et al.
Published: (2026)
by: Sijwali, Suryansh Singh, et al.
Published: (2026)
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
by: Hossain, Ismail, et al.
Published: (2026)
by: Hossain, Ismail, et al.
Published: (2026)
Detecting Quishing Attacks with Machine Learning Techniques Through QR Code Analysis
by: Trad, Fouad, et al.
Published: (2025)
by: Trad, Fouad, et al.
Published: (2025)
Vulnerabilities in AI Code Generators: Exploring Targeted Data Poisoning Attacks
by: Cotroneo, Domenico, et al.
Published: (2023)
by: Cotroneo, Domenico, et al.
Published: (2023)
Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents
by: Ding, Renhua, et al.
Published: (2025)
by: Ding, Renhua, et al.
Published: (2025)
ATAG: AI-Agent Application Threat Assessment with Attack Graphs
by: Gandhi, Parth Atulbhai, et al.
Published: (2025)
by: Gandhi, Parth Atulbhai, et al.
Published: (2025)
VibeGuard: A Security Gate Framework for AI-Generated Code
by: Xie, Ying
Published: (2026)
by: Xie, Ying
Published: (2026)
Security of AI Agents
by: He, Yifeng, et al.
Published: (2024)
by: He, Yifeng, et al.
Published: (2024)
A Systematic Evaluation of Parameter-Efficient Fine-Tuning Methods for the Security of Code LLMs
by: Lee, Kiho, et al.
Published: (2025)
by: Lee, Kiho, et al.
Published: (2025)
A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
by: Wang, Jinghao, et al.
Published: (2025)
by: Wang, Jinghao, et al.
Published: (2025)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
by: Bazinska, Julia, et al.
Published: (2025)
by: Bazinska, Julia, et al.
Published: (2025)
Jailbreaking Large Language Models through Iterative Tool-Disguised Attacks via Reinforcement Learning
by: Wang, Zhaoqi, et al.
Published: (2026)
by: Wang, Zhaoqi, et al.
Published: (2026)
CivicShield: A Cross-Domain Defense-in-Depth Framework for Securing Government-Facing AI Chatbots Against Multi-Turn Adversarial Attacks
by: Patil, KrishnaSaiReddy
Published: (2026)
by: Patil, KrishnaSaiReddy
Published: (2026)
Hallucinating AI Hijacking Attack: Large Language Models and Malicious Code Recommenders
by: Noever, David, et al.
Published: (2024)
by: Noever, David, et al.
Published: (2024)
Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success
by: Maple, Carsten, et al.
Published: (2026)
by: Maple, Carsten, et al.
Published: (2026)
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
by: Qi, Senmao, et al.
Published: (2025)
by: Qi, Senmao, et al.
Published: (2025)
Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
by: Improta, Cristina
Published: (2024)
by: Improta, Cristina
Published: (2024)
AI Code Generators for Security: Friend or Foe?
by: Natella, Roberto, et al.
Published: (2024)
by: Natella, Roberto, et al.
Published: (2024)
Security of Internet of Agents: Attacks and Countermeasures
by: Wang, Yuntao, et al.
Published: (2025)
by: Wang, Yuntao, et al.
Published: (2025)
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
MALIGN: Explainable Static Raw-byte Based Malware Family Classification using Sequence Alignment
by: Saha, Shoumik, et al.
Published: (2021)
by: Saha, Shoumik, et al.
Published: (2021)
Similar Items
-
Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
by: Saha, Shoumik, et al.
Published: (2026) -
SoK: Understanding (New) Security Issues Across AI4Code Use Cases
by: Wu, Qilong, et al.
Published: (2025) -
Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
by: Xie, Yuchong, et al.
Published: (2025) -
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024) -
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
by: Shen, Chihao, et al.
Published: (2025)