Hacking CTFs with Plain Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Turtayev, Rustem, Petrov, Artem, Volkov, Dmitrii, Volk, Denis |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events
by: Reworr, et al.
Published: (2025)
by: Reworr, et al.
Published: (2025)
Evaluating AI cyber capabilities with crowdsourced elicitation
by: Petrov, Artem, et al.
Published: (2025)
by: Petrov, Artem, et al.
Published: (2025)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
by: Reworr, et al.
Published: (2024)
by: Reworr, et al.
Published: (2024)
Autonomous LLM Agents & CTFs: A Second Look
by: Bouchari, Youness, et al.
Published: (2026)
by: Bouchari, Youness, et al.
Published: (2026)
Badllama 3: removing safety finetuning from Llama 3 in minutes
by: Volkov, Dmitrii
Published: (2024)
by: Volkov, Dmitrii
Published: (2024)
LLM Agents can Autonomously Hack Websites
by: Fang, Richard, et al.
Published: (2024)
by: Fang, Richard, et al.
Published: (2024)
Language Models Can Autonomously Hack and Self-Replicate
by: Air, Alena, et al.
Published: (2026)
by: Air, Alena, et al.
Published: (2026)
PenTest++: Elevating Ethical Hacking with AI and Automation
by: Al-Sinani, Haitham S., et al.
Published: (2025)
by: Al-Sinani, Haitham S., et al.
Published: (2025)
AI-Enhanced Ethical Hacking: A Linux-Focused Experiment
by: Al-Sinani, Haitham S., et al.
Published: (2024)
by: Al-Sinani, Haitham S., et al.
Published: (2024)
To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack
by: Zhuo, Terry Yue, et al.
Published: (2026)
by: Zhuo, Terry Yue, et al.
Published: (2026)
WiFiPenTester: Advancing Wireless Ethical Hacking with Governed GenAI
by: Al-Sinani, Haitham S., et al.
Published: (2026)
by: Al-Sinani, Haitham S., et al.
Published: (2026)
NEST: Nascent Encoded Steganographic Thoughts
by: Karpov, Artem
Published: (2026)
by: Karpov, Artem
Published: (2026)
Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks
by: Pasquini, Dario, et al.
Published: (2024)
by: Pasquini, Dario, et al.
Published: (2024)
VET Your Agent: Towards Host-Independent Autonomy via Verifiable Execution Traces
by: Grigor, Artem, et al.
Published: (2025)
by: Grigor, Artem, et al.
Published: (2025)
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
by: Schulhoff, Sander, et al.
Published: (2023)
by: Schulhoff, Sander, et al.
Published: (2023)
Misalignment Bounty: Crowdsourcing AI Agent Misbehavior
by: Turtayev, Rustem, et al.
Published: (2025)
by: Turtayev, Rustem, et al.
Published: (2025)
Hack Me If You Can: Aggregating AutoEncoders for Countering Persistent Access Threats Within Highly Imbalanced Data
by: Benabderrahmane, Sidahmed, et al.
Published: (2024)
by: Benabderrahmane, Sidahmed, et al.
Published: (2024)
Biologically-Informed Hybrid Membership Inference Attacks on Generative Genomic Models
by: Belfiore, Asia, et al.
Published: (2025)
by: Belfiore, Asia, et al.
Published: (2025)
LLM-Guided Prompt Evolution for Password Guessing
by: Mazin, Vladimir A., et al.
Published: (2026)
by: Mazin, Vladimir A., et al.
Published: (2026)
MemTrust: A Zero-Trust Architecture for Unified AI Memory System
by: Zhou, Xing, et al.
Published: (2026)
by: Zhou, Xing, et al.
Published: (2026)
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
by: Hartman, Max, et al.
Published: (2026)
by: Hartman, Max, et al.
Published: (2026)
The Hidden Threat in Plain Text: Attacking RAG Data Loaders
by: Castagnaro, Alberto, et al.
Published: (2025)
by: Castagnaro, Alberto, et al.
Published: (2025)
Hacking, The Lazy Way: LLM Augmented Pentesting
by: Goyal, Dhruva, et al.
Published: (2024)
by: Goyal, Dhruva, et al.
Published: (2024)
Cybersecurity AI: Evaluating Agentic Cybersecurity in Attack/Defense CTFs
by: Balassone, Francesco, et al.
Published: (2025)
by: Balassone, Francesco, et al.
Published: (2025)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
by: Geng, Jianing, et al.
Published: (2025)
by: Geng, Jianing, et al.
Published: (2025)
SoK: Prompt Hacking of Large Language Models
by: Rababah, Baha, et al.
Published: (2024)
by: Rababah, Baha, et al.
Published: (2024)
Incentivising the federation: gradient-based metrics for data selection and valuation in private decentralised training
by: Usynin, Dmitrii, et al.
Published: (2023)
by: Usynin, Dmitrii, et al.
Published: (2023)
Analysis of the vulnerability of machine learning regression models to adversarial attacks using data from 5G wireless networks
by: Legashev, Leonid, et al.
Published: (2025)
by: Legashev, Leonid, et al.
Published: (2025)
Agent Control Protocol: Admission Control for Agent Actions
by: Fernandez, Marcelo
Published: (2026)
by: Fernandez, Marcelo
Published: (2026)
AgentMark: Utility-Preserving Behavioral Watermarking for Agents
by: Huang, Kaibo, et al.
Published: (2026)
by: Huang, Kaibo, et al.
Published: (2026)
AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior
by: Abaev, Nadya, et al.
Published: (2026)
by: Abaev, Nadya, et al.
Published: (2026)
Agent-Sentry: Bounding LLM Agents via Execution Provenance
by: Sequeira, Rohan, et al.
Published: (2026)
by: Sequeira, Rohan, et al.
Published: (2026)
LiaisonAgent: An Multi-Agent Framework for Autonomous Risk Investigation and Governance
by: Tang, Chuanming, et al.
Published: (2026)
by: Tang, Chuanming, et al.
Published: (2026)
AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents
by: Zheng, Ye, et al.
Published: (2025)
by: Zheng, Ye, et al.
Published: (2025)
AgentWall: A Runtime Safety Layer for Local AI Agents
by: Aravind, Ashwin
Published: (2026)
by: Aravind, Ashwin
Published: (2026)
PocketAgents: A Manifest-Driven Library of Autonomous Defense Agents
by: Barbieri, Sidnei, et al.
Published: (2026)
by: Barbieri, Sidnei, et al.
Published: (2026)
Agent-Fence: Mapping Security Vulnerabilities Across Deep Research Agents
by: Puppala, Sai, et al.
Published: (2026)
by: Puppala, Sai, et al.
Published: (2026)
AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents
by: Zhang, Yixiang, et al.
Published: (2026)
by: Zhang, Yixiang, et al.
Published: (2026)
APEX: Agent Payment Execution with Policy for Autonomous Agent API Access
by: Uddin, Mohd Safwan, et al.
Published: (2026)
by: Uddin, Mohd Safwan, et al.
Published: (2026)
Agent Audit: A Security Analysis System for LLM Agent Applications
by: Zhang, Haiyue, et al.
Published: (2026)
by: Zhang, Haiyue, et al.
Published: (2026)
Similar Items
-
GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events
by: Reworr, et al.
Published: (2025) -
Evaluating AI cyber capabilities with crowdsourced elicitation
by: Petrov, Artem, et al.
Published: (2025) -
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
by: Reworr, et al.
Published: (2024) -
Autonomous LLM Agents & CTFs: A Second Look
by: Bouchari, Youness, et al.
Published: (2026) -
Badllama 3: removing safety finetuning from Llama 3 in minutes
by: Volkov, Dmitrii
Published: (2024)