Detecting Malicious AI Agents Through Simulated Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Pi, Yulu, Bettison, Ella, Becker, Anna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
by: Guo, Yihao, et al.
Published: (2025)
by: Guo, Yihao, et al.
Published: (2025)
Hierarchical Local-Global Feature Learning for Few-shot Malicious Traffic Detection
by: Peng, Songtao, et al.
Published: (2025)
by: Peng, Songtao, et al.
Published: (2025)
MULTI-LF: A Continuous Learning Framework for Real-Time Malicious Traffic Detection in Multi-Environment Networks
by: Rustam, Furqan, et al.
Published: (2025)
by: Rustam, Furqan, et al.
Published: (2025)
Enhancing GraphQL Security by Detecting Malicious Queries Using Large Language Models, Sentence Transformers, and Convolutional Neural Networks
by: Perera, Irash, et al.
Published: (2025)
by: Perera, Irash, et al.
Published: (2025)
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
by: Zloczower, Itay, et al.
Published: (2026)
by: Zloczower, Itay, et al.
Published: (2026)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
by: Bazinska, Julia, et al.
Published: (2025)
by: Bazinska, Julia, et al.
Published: (2025)
AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior
by: Abaev, Nadya, et al.
Published: (2026)
by: Abaev, Nadya, et al.
Published: (2026)
CP-Guard+: A New Paradigm for Malicious Agent Detection and Defense in Collaborative Perception
by: Hu, Senkang, et al.
Published: (2025)
by: Hu, Senkang, et al.
Published: (2025)
Detecting and Eliminating Neural Network Backdoors Through Active Paths with Application to Intrusion Detection
by: Høyheim, Eirik, et al.
Published: (2026)
by: Høyheim, Eirik, et al.
Published: (2026)
EVMbench: Evaluating AI Agents on Smart Contract Security
by: Wang, Justin, et al.
Published: (2026)
by: Wang, Justin, et al.
Published: (2026)
SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis
by: Zhang, Zhisheng, et al.
Published: (2025)
by: Zhang, Zhisheng, et al.
Published: (2025)
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
by: Hou, Yinghan, et al.
Published: (2026)
by: Hou, Yinghan, et al.
Published: (2026)
Out-of-Distribution Detection for Neurosymbolic Autonomous Cyber Agents
by: Samaddar, Ankita, et al.
Published: (2024)
by: Samaddar, Ankita, et al.
Published: (2024)
CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
by: Das, Debeshee, et al.
Published: (2025)
by: Das, Debeshee, et al.
Published: (2025)
Adversarial Reinforcement Learning for Offensive and Defensive Agents in a Simulated Zero-Sum Network Environment
by: Shahid, Abrar, et al.
Published: (2025)
by: Shahid, Abrar, et al.
Published: (2025)
Explainable AI for Comparative Analysis of Intrusion Detection Models
by: Corea, Pap M., et al.
Published: (2024)
by: Corea, Pap M., et al.
Published: (2024)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
ADR: An Agentic Detection System for Enterprise Agentic AI Security
by: Li, Chenning, et al.
Published: (2026)
by: Li, Chenning, et al.
Published: (2026)
Statement-Level Vulnerability Detection: Learning Vulnerability Patterns Through Information Theory and Contrastive Learning
by: Nguyen, Van, et al.
Published: (2022)
by: Nguyen, Van, et al.
Published: (2022)
Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
by: Guo, Qiming, et al.
Published: (2025)
by: Guo, Qiming, et al.
Published: (2025)
Undetectable Conversations Between AI Agents via Pseudorandom Noise-Resilient Key Exchange
by: Vaikuntanathan, Vinod, et al.
Published: (2026)
by: Vaikuntanathan, Vinod, et al.
Published: (2026)
How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System
by: Zuo, Kaiwen, et al.
Published: (2025)
by: Zuo, Kaiwen, et al.
Published: (2025)
Runtime Detection of Adversarial Attacks in AI Accelerators Using Performance Counters
by: Rahaman, Habibur, et al.
Published: (2025)
by: Rahaman, Habibur, et al.
Published: (2025)
ExAI5G: A Logic-Based Explainable AI Framework for Intrusion Detection in 5G Networks
by: Sheikhi, Saeid, et al.
Published: (2026)
by: Sheikhi, Saeid, et al.
Published: (2026)
Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines
by: Ahad, Tanzim, et al.
Published: (2026)
by: Ahad, Tanzim, et al.
Published: (2026)
Automated and Explainable Denial of Service Analysis for AI-Driven Intrusion Detection Systems
by: Yakubu, Paul Badu, et al.
Published: (2025)
by: Yakubu, Paul Badu, et al.
Published: (2025)
Phishing Detection in the Gen-AI Era: Quantized LLMs vs Classical Models
by: Thapa, Jikesh, et al.
Published: (2025)
by: Thapa, Jikesh, et al.
Published: (2025)
Towards Trustworthy AI: Secure Deepfake Detection using CNNs and Zero-Knowledge Proofs
by: Islam, H M Mohaimanul, et al.
Published: (2025)
by: Islam, H M Mohaimanul, et al.
Published: (2025)
Detecting Cybersecurity Threats by Integrating Explainable AI with SHAP Interpretability and Strategic Data Sampling
by: Srisumrith, Norrakith, et al.
Published: (2026)
by: Srisumrith, Norrakith, et al.
Published: (2026)
Dynamic Neural Control Flow Execution: An Agent-Based Deep Equilibrium Approach for Binary Vulnerability Detection
by: Li, Litao, et al.
Published: (2024)
by: Li, Litao, et al.
Published: (2024)
Improving IoT Intrusion Detection Through SMOTE-Based Oversampling and Extended Multi-Model Evaluation on Side-Channel Power Data
by: Shahzad, Muhammad Khuram, et al.
Published: (2026)
by: Shahzad, Muhammad Khuram, et al.
Published: (2026)
Secure Energy Transactions Using Blockchain Leveraging AI for Fraud Detection and Energy Market Stability
by: Khan, Md Asif Ul Hoq, et al.
Published: (2025)
by: Khan, Md Asif Ul Hoq, et al.
Published: (2025)
Exploring Feature Importance and Explainability Towards Enhanced ML-Based DoS Detection in AI Systems
by: Yakubu, Paul Badu, et al.
Published: (2024)
by: Yakubu, Paul Badu, et al.
Published: (2024)
Content-Aware Attack Detection in LLM Agent Tool-Call Traffic: An Empirical Study of Features, Architectures, and Evaluation Protocols
by: Zavrak, Sultan
Published: (2026)
by: Zavrak, Sultan
Published: (2026)
Scalable Hierarchical AI-Blockchain Framework for Real-Time Anomaly Detection in Large-Scale Autonomous Vehicle Networks
by: Shit, Rathin Chandra, et al.
Published: (2025)
by: Shit, Rathin Chandra, et al.
Published: (2025)
Quantum-Augmented AI/ML for O-RAN: Hierarchical Threat Detection with Synergistic Intelligence and Interpretability (Technical Report)
by: Le, Tan, et al.
Published: (2025)
by: Le, Tan, et al.
Published: (2025)
XAI-SOH-FL: Enhancing SOH-FL with Adaptive Aggregation and Explainable AI for Intrusion Detection in Heterogeneous IoT
by: Aslam, Ambreen, et al.
Published: (2026)
by: Aslam, Ambreen, et al.
Published: (2026)
Similar Items
-
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
by: Guo, Yihao, et al.
Published: (2025) -
Hierarchical Local-Global Feature Learning for Few-shot Malicious Traffic Detection
by: Peng, Songtao, et al.
Published: (2025) -
MULTI-LF: A Continuous Learning Framework for Real-Time Malicious Traffic Detection in Multi-Environment Networks
by: Rustam, Furqan, et al.
Published: (2025) -
Enhancing GraphQL Security by Detecting Malicious Queries Using Large Language Models, Sentence Transformers, and Convolutional Neural Networks
by: Perera, Irash, et al.
Published: (2025) -
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
by: Zloczower, Itay, et al.
Published: (2026)