A Framework for Evaluating Emerging Cyberattack Capabilities of AI
Fuente:
arXiv
Saved in:
| Main Authors: | Rodriguez, Mikel, Popa, Raluca Ada, Flynn, Four, Liang, Lihao, Dafoe, Allan, Wang, Anna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Opal: Private Memory for Personal AI
by: Kaviani, Darya, et al.
Published: (2026)
by: Kaviani, Darya, et al.
Published: (2026)
MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents
by: Zhu, Jinhao, et al.
Published: (2025)
by: Zhu, Jinhao, et al.
Published: (2025)
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
by: Lu, Yu-An, et al.
Published: (2026)
by: Lu, Yu-An, et al.
Published: (2026)
Onyx: Cost-Efficient Disk-Oblivious ANN Search
by: Rathee, Deevashwer, et al.
Published: (2026)
by: Rathee, Deevashwer, et al.
Published: (2026)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
Tackling Cyberattacks through AI-based Reactive Systems: A Holistic Review and Future Vision
by: Molina, Sergio Bernardez, et al.
Published: (2023)
by: Molina, Sergio Bernardez, et al.
Published: (2023)
The Role and Applications of Airport Digital Twin in Cyberattack Protection during the Generative AI Era
by: Weinberg, Abraham Itzhak
Published: (2024)
by: Weinberg, Abraham Itzhak
Published: (2024)
Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks
by: Pasquini, Dario, et al.
Published: (2024)
by: Pasquini, Dario, et al.
Published: (2024)
Binary and Multiclass Cyberattack Classification on GeNIS Dataset
by: Silva, Miguel, et al.
Published: (2025)
by: Silva, Miguel, et al.
Published: (2025)
Multi-Granular Discretization for Interpretable Generalization in Precise Cyberattack Identification
by: Chung, Wen-Cheng, et al.
Published: (2025)
by: Chung, Wen-Cheng, et al.
Published: (2025)
MPC-Minimized Secure LLM Inference
by: Rathee, Deevashwer, et al.
Published: (2024)
by: Rathee, Deevashwer, et al.
Published: (2024)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
by: Zhu, Xiaoyuan, et al.
Published: (2025)
by: Zhu, Xiaoyuan, et al.
Published: (2025)
The Best Defense is a Good Offense: Countering LLM-Powered Cyberattacks
by: Ayzenshteyn, Daniel, et al.
Published: (2024)
by: Ayzenshteyn, Daniel, et al.
Published: (2024)
Web Agents Should Adopt the Plan-Then-Execute Paradigm
by: Piet, Julien, et al.
Published: (2026)
by: Piet, Julien, et al.
Published: (2026)
Large Language Model-Based Framework for Explainable Cyberattack Detection in Automatic Generation Control Systems
by: Sharshar, Muhammad, et al.
Published: (2025)
by: Sharshar, Muhammad, et al.
Published: (2025)
Simulating Cyberattacks through a Breach Attack Simulation (BAS) Platform empowered by Security Chaos Engineering (SCE)
by: Sánchez-Matas, Arturo, et al.
Published: (2025)
by: Sánchez-Matas, Arturo, et al.
Published: (2025)
Emerging Cyber Attack Risks of Medical AI Agents
by: Qiu, Jianing, et al.
Published: (2025)
by: Qiu, Jianing, et al.
Published: (2025)
AI-Driven Cybersecurity Threats: A Survey of Emerging Risks and Defensive Strategies
by: Erukude, Sai Teja, et al.
Published: (2026)
by: Erukude, Sai Teja, et al.
Published: (2026)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
by: Wang, Jinghao, et al.
Published: (2025)
by: Wang, Jinghao, et al.
Published: (2025)
Twin Auto-Encoder Model for Learning Separable Representation in Cyberattack Detection
by: Dinh, Phai Vu, et al.
Published: (2024)
by: Dinh, Phai Vu, et al.
Published: (2024)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)
by: Che, Zora, et al.
Published: (2025)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
by: Black, Sid, et al.
Published: (2025)
by: Black, Sid, et al.
Published: (2025)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
by: Kouremetis, Michael, et al.
Published: (2025)
by: Kouremetis, Michael, et al.
Published: (2025)
Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Models
by: Ganiuly, Daniyal, et al.
Published: (2025)
by: Ganiuly, Daniyal, et al.
Published: (2025)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
by: Li, Chloe, et al.
Published: (2025)
by: Li, Chloe, et al.
Published: (2025)
AIAuditTrack: A Framework for AI Security system
by: Luo, Zixun, et al.
Published: (2025)
by: Luo, Zixun, et al.
Published: (2025)
Security Threat Modeling for Emerging AI-Agent Protocols: A Comparative Analysis of MCP, A2A, Agora, and ANP
by: Anbiaee, Zeynab, et al.
Published: (2026)
by: Anbiaee, Zeynab, et al.
Published: (2026)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
by: Ma, Haokai, et al.
Published: (2025)
by: Ma, Haokai, et al.
Published: (2025)
Operationalising Cyber Risk Management Using AI: Connecting Cyber Incidents to MITRE ATT&CK Techniques, Security Controls, and Metrics
by: Sherif, Emad, et al.
Published: (2026)
by: Sherif, Emad, et al.
Published: (2026)
STRIDE-AI: A Threat Modeling Framework for Generative AI Security Assessment
by: Cyrille, Tsafac Nkombong Regine, et al.
Published: (2026)
by: Cyrille, Tsafac Nkombong Regine, et al.
Published: (2026)
Privacy in Responsible AI: Approaches to Facial Recognition from Cloud Providers
by: Elivanova, Anna
Published: (2025)
by: Elivanova, Anna
Published: (2025)
A Comparative Evaluation of AI Agent Security Guardrails
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents
by: Narajala, Vineeth Sai, et al.
Published: (2025)
by: Narajala, Vineeth Sai, et al.
Published: (2025)
NetMoniAI: An Agentic AI Framework for Network Security & Monitoring
by: Zambare, Pallavi, et al.
Published: (2025)
by: Zambare, Pallavi, et al.
Published: (2025)
SAGE: A Generic Framework for LLM Safety Evaluation
by: Jindal, Madhur, et al.
Published: (2025)
by: Jindal, Madhur, et al.
Published: (2025)
A Framework for Cryptographic Verifiability of End-to-End AI Pipelines
by: Balan, Kar, et al.
Published: (2025)
by: Balan, Kar, et al.
Published: (2025)
A Security Analysis of the OpenClaw AI Agent Framework
by: Suwansathit, Surada, et al.
Published: (2026)
by: Suwansathit, Surada, et al.
Published: (2026)
Towards Secure and Private AI: A Framework for Decentralized Inference
by: Zhang, Hongyang, et al.
Published: (2024)
by: Zhang, Hongyang, et al.
Published: (2024)
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
Similar Items
-
Opal: Private Memory for Personal AI
by: Kaviani, Darya, et al.
Published: (2026) -
MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents
by: Zhu, Jinhao, et al.
Published: (2025) -
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
by: Lu, Yu-An, et al.
Published: (2026) -
Onyx: Cost-Efficient Disk-Oblivious ANN Search
by: Rathee, Deevashwer, et al.
Published: (2026) -
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)