ACDZero: MCTS Agent for Mastering Automated Cyber Defense

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Yu, Tang, Sizhe, Chen, Rongqian, Yu, Fei Xu, Jiang, Guangyu, Imani, Mahdi, Bastian, Nathaniel D., Lan, Tian
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915719231307776
author Li, Yu
Tang, Sizhe
Chen, Rongqian
Yu, Fei Xu
Jiang, Guangyu
Imani, Mahdi
Bastian, Nathaniel D.
Lan, Tian
author_facet Li, Yu
Tang, Sizhe
Chen, Rongqian
Yu, Fei Xu
Jiang, Guangyu
Imani, Mahdi
Bastian, Nathaniel D.
Lan, Tian
contents Automated cyber defense (ACD) seeks to protect computer networks with minimal or no human intervention, reacting to intrusions by taking corrective actions such as isolating hosts, resetting services, deploying decoys, or updating access controls. However, existing approaches for ACD, such as deep reinforcement learning (RL), often face difficult exploration in complex networks with large decision/state spaces and thus require an expensive amount of samples. Inspired by the need to learn sample-efficient defense policies, we frame ACD in CAGE Challenge 4 (CAGE-4 / CC4) as a context-based partially observable Markov decision problem and propose a planning-centric defense policy based on Monte Carlo Tree Search (MCTS). It explicitly models the exploration-exploitation tradeoff in ACD and uses statistical sampling to guide exploration and decision making. We make novel use of graph neural networks (GNNs) to embed observations from the network as attributed graphs, to enable permutation-invariant reasoning over hosts and their relationships. To make our solution practical in complex search spaces, we guide MCTS with learned graph embeddings and priors over graph-edit actions, combining model-free generalization and policy distillation with look-ahead planning. We evaluate the resulting agent on CC4 scenarios involving diverse network structures and adversary behaviors, and show that our search-guided, graph-embedding-based planning improves defense reward and robustness relative to state-of-the-art RL baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02196
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ACDZero: MCTS Agent for Mastering Automated Cyber Defense
Li, Yu
Tang, Sizhe
Chen, Rongqian
Yu, Fei Xu
Jiang, Guangyu
Imani, Mahdi
Bastian, Nathaniel D.
Lan, Tian
Machine Learning
Automated cyber defense (ACD) seeks to protect computer networks with minimal or no human intervention, reacting to intrusions by taking corrective actions such as isolating hosts, resetting services, deploying decoys, or updating access controls. However, existing approaches for ACD, such as deep reinforcement learning (RL), often face difficult exploration in complex networks with large decision/state spaces and thus require an expensive amount of samples. Inspired by the need to learn sample-efficient defense policies, we frame ACD in CAGE Challenge 4 (CAGE-4 / CC4) as a context-based partially observable Markov decision problem and propose a planning-centric defense policy based on Monte Carlo Tree Search (MCTS). It explicitly models the exploration-exploitation tradeoff in ACD and uses statistical sampling to guide exploration and decision making. We make novel use of graph neural networks (GNNs) to embed observations from the network as attributed graphs, to enable permutation-invariant reasoning over hosts and their relationships. To make our solution practical in complex search spaces, we guide MCTS with learned graph embeddings and priors over graph-edit actions, combining model-free generalization and policy distillation with look-ahead planning. We evaluate the resulting agent on CC4 scenarios involving diverse network structures and adversary behaviors, and show that our search-guided, graph-embedding-based planning improves defense reward and robustness relative to state-of-the-art RL baselines.
title ACDZero: MCTS Agent for Mastering Automated Cyber Defense
topic Machine Learning
url https://arxiv.org/abs/2601.02196