LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
Fuente:
arXiv
Saved in:
| Main Authors: | Reworr, Volkov, Dmitrii |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hacking CTFs with Plain Agents
by: Turtayev, Rustem, et al.
Published: (2024)
by: Turtayev, Rustem, et al.
Published: (2024)
Language Models Can Autonomously Hack and Self-Replicate
by: Air, Alena, et al.
Published: (2026)
by: Air, Alena, et al.
Published: (2026)
GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events
by: Reworr, et al.
Published: (2025)
by: Reworr, et al.
Published: (2025)
LLM Agents can Autonomously Hack Websites
by: Fang, Richard, et al.
Published: (2024)
by: Fang, Richard, et al.
Published: (2024)
Evaluating AI cyber capabilities with crowdsourced elicitation
by: Petrov, Artem, et al.
Published: (2025)
by: Petrov, Artem, et al.
Published: (2025)
Design and Development of an Intelligent LLM-based LDAP Honeypot
by: Jiménez-Román, Javier, et al.
Published: (2025)
by: Jiménez-Román, Javier, et al.
Published: (2025)
LLM in the Shell: Generative Honeypots
by: Sladić, Muris, et al.
Published: (2023)
by: Sladić, Muris, et al.
Published: (2023)
HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
Badllama 3: removing safety finetuning from Llama 3 in minutes
by: Volkov, Dmitrii
Published: (2024)
by: Volkov, Dmitrii
Published: (2024)
To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack
by: Zhuo, Terry Yue, et al.
Published: (2026)
by: Zhuo, Terry Yue, et al.
Published: (2026)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
by: Wu, ChenYu, et al.
Published: (2025)
by: Wu, ChenYu, et al.
Published: (2025)
Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks
by: Pasquini, Dario, et al.
Published: (2024)
by: Pasquini, Dario, et al.
Published: (2024)
PenTest++: Elevating Ethical Hacking with AI and Automation
by: Al-Sinani, Haitham S., et al.
Published: (2025)
by: Al-Sinani, Haitham S., et al.
Published: (2025)
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
by: Conde, Pedro, et al.
Published: (2026)
by: Conde, Pedro, et al.
Published: (2026)
An Application-Layer Multi-Modal Covert-Channel Reference Monitor for LLM Agent Egress
by: Metere, Alfredo
Published: (2026)
by: Metere, Alfredo
Published: (2026)
Reliable Weak-to-Strong Monitoring of LLM Agents
by: Kale, Neil, et al.
Published: (2025)
by: Kale, Neil, et al.
Published: (2025)
AI-Enhanced Ethical Hacking: A Linux-Focused Experiment
by: Al-Sinani, Haitham S., et al.
Published: (2024)
by: Al-Sinani, Haitham S., et al.
Published: (2024)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
by: Jotautaitė, Monika, et al.
Published: (2026)
by: Jotautaitė, Monika, et al.
Published: (2026)
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
by: Xiang, Qiuchi, et al.
Published: (2026)
by: Xiang, Qiuchi, et al.
Published: (2026)
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
by: Vero, Mark, et al.
Published: (2026)
by: Vero, Mark, et al.
Published: (2026)
CheatAgent: Attacking LLM-Empowered Recommender Systems via LLM Agent
by: Ning, Liang-bo, et al.
Published: (2025)
by: Ning, Liang-bo, et al.
Published: (2025)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Security of AI Agents
by: He, Yifeng, et al.
Published: (2024)
by: He, Yifeng, et al.
Published: (2024)
Agent-Sentry: Bounding LLM Agents via Execution Provenance
by: Sequeira, Rohan, et al.
Published: (2026)
by: Sequeira, Rohan, et al.
Published: (2026)
Agent Audit: A Security Analysis System for LLM Agent Applications
by: Zhang, Haiyue, et al.
Published: (2026)
by: Zhang, Haiyue, et al.
Published: (2026)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
Securing LLM-Generated Embedded Firmware through AI Agent-Driven Validation and Patching
by: Abtahi, Seyed Moein, et al.
Published: (2025)
by: Abtahi, Seyed Moein, et al.
Published: (2025)
WiFiPenTester: Advancing Wireless Ethical Hacking with Governed GenAI
by: Al-Sinani, Haitham S., et al.
Published: (2026)
by: Al-Sinani, Haitham S., et al.
Published: (2026)
AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents
by: Zheng, Ye, et al.
Published: (2025)
by: Zheng, Ye, et al.
Published: (2025)
AgentWall: A Runtime Safety Layer for Local AI Agents
by: Aravind, Ashwin
Published: (2026)
by: Aravind, Ashwin
Published: (2026)
AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents
by: Zhang, Yixiang, et al.
Published: (2026)
by: Zhang, Yixiang, et al.
Published: (2026)
Sequential Behavioral Watermarking for LLM Agents
by: An, Hyeseon, et al.
Published: (2026)
by: An, Hyeseon, et al.
Published: (2026)
Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms
by: Gasmi, Tarek, et al.
Published: (2025)
by: Gasmi, Tarek, et al.
Published: (2025)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent
by: Zhu, Pengyu, et al.
Published: (2025)
by: Zhu, Pengyu, et al.
Published: (2025)
AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management
by: Wen, Ruoyao, et al.
Published: (2026)
by: Wen, Ruoyao, et al.
Published: (2026)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
by: Yang, Chenglin
Published: (2026)
by: Yang, Chenglin
Published: (2026)
Reasoning-Style Poisoning of LLM Agents via Stealthy Style Transfer: Process-Level Attacks and Runtime Monitoring in RSV Space
by: Zhou, Xingfu, et al.
Published: (2025)
by: Zhou, Xingfu, et al.
Published: (2025)
The Hidden Dangers of Browsing AI Agents
by: Mudryi, Mykyta, et al.
Published: (2025)
by: Mudryi, Mykyta, et al.
Published: (2025)
Similar Items
-
Hacking CTFs with Plain Agents
by: Turtayev, Rustem, et al.
Published: (2024) -
Language Models Can Autonomously Hack and Self-Replicate
by: Air, Alena, et al.
Published: (2026) -
GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events
by: Reworr, et al.
Published: (2025) -
LLM Agents can Autonomously Hack Websites
by: Fang, Richard, et al.
Published: (2024) -
Evaluating AI cyber capabilities with crowdsourced elicitation
by: Petrov, Artem, et al.
Published: (2025)