Saved in:
Bibliographic Details
Main Authors: Hugglestone, James, Chacko, Samuel Jacob, Stoller, Dawson, Schmidt, Ryan, Liu, Xiuwen
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.22577
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908908690341888
author Hugglestone, James
Chacko, Samuel Jacob
Stoller, Dawson
Schmidt, Ryan
Liu, Xiuwen
author_facet Hugglestone, James
Chacko, Samuel Jacob
Stoller, Dawson
Schmidt, Ryan
Liu, Xiuwen
contents Large Language Models (LLMs) have demonstrated potential in code generation, yet they struggle with the multi-step, stateful reasoning required for offensive cybersecurity operations. Existing research often relies on static benchmarks that fail to capture the dynamic nature of real-world vulnerabilities. In this work, we introduce STRIATUM-CTF (A Search-based Test-time Reasoning Inference Agent for Tactical Utility Maximization in Cybersecurity), a modular agentic framework built upon the Model Context Protocol (MCP). By standardizing tool interfaces for system introspection, decompilation, and runtime debugging, STRIATUM-CTF enables the agent to maintain a coherent context window across extended exploit trajectories. We validate this approach not merely on synthetic datasets, but in a live competitive environment. Our system participated in a university-hosted Capture-the-Flag (CTF) competition in late 2025, where it operated autonomously to identify and exploit vulnerabilities in real-time. STRIATUM-CTF secured First Place, outperforming 21 human teams and demonstrating strong adaptability in a dynamic problem-solving setting. We analyze the agent's decision-making logs to show how MCP-based tool abstraction significantly reduces hallucination compared to naive prompting strategies. These results suggest that standardized context protocols are a critical path toward robust autonomous cyber-reasoning systems.
format Preprint
id arxiv_https___arxiv_org_abs_2603_22577
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STRIATUM-CTF: A Protocol-Driven Agentic Framework for General-Purpose CTF Solving
Hugglestone, James
Chacko, Samuel Jacob
Stoller, Dawson
Schmidt, Ryan
Liu, Xiuwen
Cryptography and Security
Artificial Intelligence
Multiagent Systems
I.2.1
Large Language Models (LLMs) have demonstrated potential in code generation, yet they struggle with the multi-step, stateful reasoning required for offensive cybersecurity operations. Existing research often relies on static benchmarks that fail to capture the dynamic nature of real-world vulnerabilities. In this work, we introduce STRIATUM-CTF (A Search-based Test-time Reasoning Inference Agent for Tactical Utility Maximization in Cybersecurity), a modular agentic framework built upon the Model Context Protocol (MCP). By standardizing tool interfaces for system introspection, decompilation, and runtime debugging, STRIATUM-CTF enables the agent to maintain a coherent context window across extended exploit trajectories. We validate this approach not merely on synthetic datasets, but in a live competitive environment. Our system participated in a university-hosted Capture-the-Flag (CTF) competition in late 2025, where it operated autonomously to identify and exploit vulnerabilities in real-time. STRIATUM-CTF secured First Place, outperforming 21 human teams and demonstrating strong adaptability in a dynamic problem-solving setting. We analyze the agent's decision-making logs to show how MCP-based tool abstraction significantly reduces hallucination compared to naive prompting strategies. These results suggest that standardized context protocols are a critical path toward robust autonomous cyber-reasoning systems.
title STRIATUM-CTF: A Protocol-Driven Agentic Framework for General-Purpose CTF Solving
topic Cryptography and Security
Artificial Intelligence
Multiagent Systems
I.2.1
url https://arxiv.org/abs/2603.22577