The Automation Advantage in AI Red Teaming
Fuente:
arXiv
Saved in:
| Main Authors: | Mulla, Rob, Dawson, Ads, Abruzzon, Vincent, Greunke, Brian, Landers, Nick, Palm, Brad, Pearce, Will |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
by: Dawson, Ads, et al.
Published: (2025)
by: Dawson, Ads, et al.
Published: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
by: Othman, Refat
Published: (2026)
by: Othman, Refat
Published: (2026)
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
by: Grofsky, Matthew
Published: (2025)
by: Grofsky, Matthew
Published: (2025)
Multilingual AI-Driven Password Strength Estimation with Similarity-Based Detection
by: Palaniappan, Nikitha M., et al.
Published: (2026)
by: Palaniappan, Nikitha M., et al.
Published: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study
by: Khatiwala, Jeel Piyushkumar, et al.
Published: (2026)
by: Khatiwala, Jeel Piyushkumar, et al.
Published: (2026)
Towards Agentic Investigation of Security Alerts
by: Eilertsen, Even, et al.
Published: (2026)
by: Eilertsen, Even, et al.
Published: (2026)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
by: Zhang, Tian, et al.
Published: (2026)
by: Zhang, Tian, et al.
Published: (2026)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
by: Ge, Yuxu
Published: (2026)
by: Ge, Yuxu
Published: (2026)
Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
by: Adebimpe, Adetayo, et al.
Published: (2025)
by: Adebimpe, Adetayo, et al.
Published: (2025)
Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
by: Collu, Matteo Gioele, et al.
Published: (2023)
by: Collu, Matteo Gioele, et al.
Published: (2023)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
by: Nathanson, Samuel, et al.
Published: (2025)
by: Nathanson, Samuel, et al.
Published: (2025)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
by: Gu, Yongtong, et al.
Published: (2026)
by: Gu, Yongtong, et al.
Published: (2026)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
by: Chakraborty, Amit, et al.
Published: (2025)
by: Chakraborty, Amit, et al.
Published: (2025)
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
by: Morales, Jaime, et al.
Published: (2026)
by: Morales, Jaime, et al.
Published: (2026)
Static Attribution of Android Residential Proxy Malware Using Graph Kernels
by: Clark, Peter, et al.
Published: (2026)
by: Clark, Peter, et al.
Published: (2026)
Security Considerations for Multi-agent Systems
by: Nguyen, Tam, et al.
Published: (2026)
by: Nguyen, Tam, et al.
Published: (2026)
Memory Forensics Techniques for Automated Detection and Analysis of Go Malware
by: Ali, Hala, et al.
Published: (2026)
by: Ali, Hala, et al.
Published: (2026)
ILION: Deterministic Pre-Execution Safety Gates for Agentic AI Systems
by: Chitan, Florin Adrian
Published: (2026)
by: Chitan, Florin Adrian
Published: (2026)
Hiding in the AI Traffic: Abusing MCP for LLM-Powered Agentic Red Teaming
by: Janjusevic, Strahinja, et al.
Published: (2025)
by: Janjusevic, Strahinja, et al.
Published: (2025)
Safeguarding Virtual Healthcare: A Novel Attacker-Centric Model for Data Security and Privacy
by: Herath, Suvineetha, et al.
Published: (2024)
by: Herath, Suvineetha, et al.
Published: (2024)
A Secure, Manifest-Based Framework for Delegated Privilege Promotion
by: Chowdhury, Rajarshi, et al.
Published: (2026)
by: Chowdhury, Rajarshi, et al.
Published: (2026)
Can AI Keep a Secret? Contextual Integrity Verification: A Provable Security Architecture for LLMs
by: Gupta, Aayush
Published: (2025)
by: Gupta, Aayush
Published: (2025)
Session Risk Memory (SRM): Temporal Authorization for Deterministic Pre-Execution Safety Gates
by: Chitan, Florin Adrian
Published: (2026)
by: Chitan, Florin Adrian
Published: (2026)
Explainable Attention-Based LSTM Framework for Early Detection of AI-Assisted Ransomware via File System Behavioral Analysis
by: Nayak, Prabhudarshi, et al.
Published: (2026)
by: Nayak, Prabhudarshi, et al.
Published: (2026)
Send to which account? Evaluation of an LLM-based Scambaiting System
by: Siadati, Hossein, et al.
Published: (2025)
by: Siadati, Hossein, et al.
Published: (2025)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
by: Chona, Alankrit, et al.
Published: (2026)
by: Chona, Alankrit, et al.
Published: (2026)
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Measuring Harmfulness of Computer-Using Agents
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
Countermind: A Multi-Layered Security Architecture for Large Language Models
by: Schwarz, Dominik
Published: (2025)
by: Schwarz, Dominik
Published: (2025)
ARMOUR US: Android Runtime Zero-permission Sensor Usage Monitoring from User Space
by: Long, Yan, et al.
Published: (2025)
by: Long, Yan, et al.
Published: (2025)
The Discovery, Disclosure, and Investigation of CVE-2024-25825
by: Chasens, Hunter
Published: (2025)
by: Chasens, Hunter
Published: (2025)
Post-quantum Federated Learning: Secure And Scalable Threat Intelligence For Collaborative Cyber Defense
by: Nayak, Prabhudarshi, et al.
Published: (2026)
by: Nayak, Prabhudarshi, et al.
Published: (2026)
Silent Consent, Persistent Risk: Android Permission Groups and Custom Permissions
by: Akanji, Olawale Amos, et al.
Published: (2026)
by: Akanji, Olawale Amos, et al.
Published: (2026)
Dynamic Vulnerability Patching for Heterogeneous Embedded Systems Using Stack Frame Reconstruction
by: Zhou, Ming, et al.
Published: (2025)
by: Zhou, Ming, et al.
Published: (2025)
vCause: Efficient and Verifiable Causality Analysis for Cloud-based Endpoint Auditing
by: Song, Qiyang, et al.
Published: (2026)
by: Song, Qiyang, et al.
Published: (2026)
Securing Agentic AI Systems -- A Multilayer Security Framework
by: Arora, Sunil, et al.
Published: (2025)
by: Arora, Sunil, et al.
Published: (2025)
Right to History: A Sovereignty Kernel for Verifiable AI Agent Execution
by: Zhang, Jing
Published: (2026)
by: Zhang, Jing
Published: (2026)
SALLIE: Safeguarding Against Latent Language & Image Exploits
by: Azov, Guy, et al.
Published: (2026)
by: Azov, Guy, et al.
Published: (2026)
Similar Items
-
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
by: Dawson, Ads, et al.
Published: (2025) -
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
by: Othman, Refat
Published: (2026) -
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
by: Grofsky, Matthew
Published: (2025) -
Multilingual AI-Driven Password Strength Estimation with Similarity-Based Detection
by: Palaniappan, Nikitha M., et al.
Published: (2026) -
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)