To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack
Fuente:
arXiv
Saved in:
| Main Authors: | Zhuo, Terry Yue, Ding, Yangruibo, Guo, Wenbo, Meng, Ruijie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Impact of AI on the Cyber Offense-Defense Balance and the Character of Cyber Conflict
by: Lohn, Andrew J.
Published: (2025)
by: Lohn, Andrew J.
Published: (2025)
Accelerating AI Development with Cyber Arenas
by: Cashman, William, et al.
Published: (2025)
by: Cashman, William, et al.
Published: (2025)
AI-Driven Cyber Threat Intelligence Automation
by: Shah, Shrit, et al.
Published: (2024)
by: Shah, Shrit, et al.
Published: (2024)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
by: Zhu, Kaijie, et al.
Published: (2025)
by: Zhu, Kaijie, et al.
Published: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
Countering Autonomous Cyber Threats
by: Heckel, Kade M., et al.
Published: (2024)
by: Heckel, Kade M., et al.
Published: (2024)
Frontier AI's Impact on the Cybersecurity Landscape
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
Artificial Intelligence in Cybersecurity: Building Resilient Cyber Diplomacy Frameworks
by: Stoltz, Michael
Published: (2024)
by: Stoltz, Michael
Published: (2024)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Emerging Cyber Attack Risks of Medical AI Agents
by: Qiu, Jianing, et al.
Published: (2025)
by: Qiu, Jianing, et al.
Published: (2025)
No Free Lunch for Defending Against Prefilling Attack by In-Context Learning
by: Xue, Zhiyu, et al.
Published: (2024)
by: Xue, Zhiyu, et al.
Published: (2024)
Defending Against Beta Poisoning Attacks in Machine Learning Models
by: Gulciftci, Nilufer, et al.
Published: (2025)
by: Gulciftci, Nilufer, et al.
Published: (2025)
Combining Threat Intelligence with IoT Scanning to Predict Cyber Attack
by: Soni, Jubin Abhishek, et al.
Published: (2024)
by: Soni, Jubin Abhishek, et al.
Published: (2024)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
by: Lin, Justin W., et al.
Published: (2025)
by: Lin, Justin W., et al.
Published: (2025)
Defending LVLMs Against Vision Attacks through Partial-Perception Supervision
by: Zhou, Qi, et al.
Published: (2024)
by: Zhou, Qi, et al.
Published: (2024)
Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
by: Xu, Wenpeng
Published: (2026)
by: Xu, Wenpeng
Published: (2026)
MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots
by: Wang, Yuhao, et al.
Published: (2026)
by: Wang, Yuhao, et al.
Published: (2026)
Inferring Discussion Topics about Exploitation of Vulnerabilities from Underground Hacking Forums
by: Moreno-Vera, Felipe
Published: (2024)
by: Moreno-Vera, Felipe
Published: (2024)
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
by: Liu, Tiantian, et al.
Published: (2024)
by: Liu, Tiantian, et al.
Published: (2024)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
by: Reworr, et al.
Published: (2024)
by: Reworr, et al.
Published: (2024)
A Whole New World: Creating a Parallel-Poisoned Web Only AI-Agents Can See
by: Zychlinski, Shaked
Published: (2025)
by: Zychlinski, Shaked
Published: (2025)
Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks
by: Jin, Haotian, et al.
Published: (2025)
by: Jin, Haotian, et al.
Published: (2025)
Complete Evasion, Zero Modification: PDF Attacks on AI Text Detection
by: Creo, Aldan
Published: (2025)
by: Creo, Aldan
Published: (2025)
Securing AI Agents Against Prompt Injection Attacks
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
AI-Powered Spearphishing Cyber Attacks: Fact or Fiction?
by: Kemp, Matthew, et al.
Published: (2025)
by: Kemp, Matthew, et al.
Published: (2025)
Co-PatcheR: Collaborative Software Patching with Component(s)-specific Small Reasoning Models
by: Tang, Yuheng, et al.
Published: (2025)
by: Tang, Yuheng, et al.
Published: (2025)
Defending Against Intelligent Attackers at Large Scales
by: Lohn, Andrew J.
Published: (2025)
by: Lohn, Andrew J.
Published: (2025)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
by: Zhong, Peter Yong, et al.
Published: (2025)
by: Zhong, Peter Yong, et al.
Published: (2025)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023)
by: Cao, Bochuan, et al.
Published: (2023)
Cyber Shadows: Neutralizing Security Threats with AI and Targeted Policy Measures
by: Schmitt, Marc, et al.
Published: (2025)
by: Schmitt, Marc, et al.
Published: (2025)
Parallax: Why AI Agents That Think Must Never Act
by: Fokou, Joel
Published: (2026)
by: Fokou, Joel
Published: (2026)
AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
by: Lin, Zhiqiang, et al.
Published: (2025)
by: Lin, Zhiqiang, et al.
Published: (2025)
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
by: Barrett, Anthony M., et al.
Published: (2025)
by: Barrett, Anthony M., et al.
Published: (2025)
AI Agents Under EU Law
by: Nannini, Luca, et al.
Published: (2026)
by: Nannini, Luca, et al.
Published: (2026)
Asymmetry by Design: Boosting Cyber Defenders with Differential Access to AI
by: Ee, Shaun, et al.
Published: (2025)
by: Ee, Shaun, et al.
Published: (2025)
Hacking CTFs with Plain Agents
by: Turtayev, Rustem, et al.
Published: (2024)
by: Turtayev, Rustem, et al.
Published: (2024)
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
by: Tong, Haibo, et al.
Published: (2026)
by: Tong, Haibo, et al.
Published: (2026)
Comparative Survey of Cyber-Threat and Attack Trends and Prediction of Future Cyber-Attack Patterns
by: Chinanu, Uwazie Emmanuel, et al.
Published: (2024)
by: Chinanu, Uwazie Emmanuel, et al.
Published: (2024)
Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
by: Wong, Ryan, et al.
Published: (2025)
by: Wong, Ryan, et al.
Published: (2025)
Similar Items
-
The Impact of AI on the Cyber Offense-Defense Balance and the Character of Cyber Conflict
by: Lohn, Andrew J.
Published: (2025) -
Accelerating AI Development with Cyber Arenas
by: Cashman, William, et al.
Published: (2025) -
AI-Driven Cyber Threat Intelligence Automation
by: Shah, Shrit, et al.
Published: (2024) -
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
by: Zhu, Kaijie, et al.
Published: (2025) -
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)