Towards Production-Worthy Simulation for Autonomous Cyber Operations

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tholl, Konur, Mezouar, Mariam El, Taylor, Adrian, Mallah, Ranwa Al
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908833978253312
author Tholl, Konur
Mezouar, Mariam El
Taylor, Adrian
Mallah, Ranwa Al
author_facet Tholl, Konur
Mezouar, Mariam El
Taylor, Adrian
Mallah, Ranwa Al
contents Simulated environments have proven invaluable in Autonomous Cyber Operations (ACO) where Reinforcement Learning (RL) agents can be trained without the computational overhead of emulation. These environments must accurately represent cybersecurity scenarios while producing the necessary signals to support RL training. In this study, we present a framework where we first extend CybORG's Cage Challenge 2 environment by implementing three new actions: Patch, Isolate, and Unisolate, to better represent the capabilities available to human operators in real-world settings. We then propose a design for agent development where we modify the reward signals and the agent's feature space to enhance training performance. To validate these modifications, we train DQN and PPO agents in the updated environment. Our study demonstrates that CybORG can be extended with additional realistic functionality, while maintaining its ability to generate informative training signals for RL agents.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19278
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Production-Worthy Simulation for Autonomous Cyber Operations
Tholl, Konur
Mezouar, Mariam El
Taylor, Adrian
Mallah, Ranwa Al
Cryptography and Security
Artificial Intelligence
Machine Learning
Simulated environments have proven invaluable in Autonomous Cyber Operations (ACO) where Reinforcement Learning (RL) agents can be trained without the computational overhead of emulation. These environments must accurately represent cybersecurity scenarios while producing the necessary signals to support RL training. In this study, we present a framework where we first extend CybORG's Cage Challenge 2 environment by implementing three new actions: Patch, Isolate, and Unisolate, to better represent the capabilities available to human operators in real-world settings. We then propose a design for agent development where we modify the reward signals and the agent's feature space to enhance training performance. To validate these modifications, we train DQN and PPO agents in the updated environment. Our study demonstrates that CybORG can be extended with additional realistic functionality, while maintaining its ability to generate informative training signals for RL agents.
title Towards Production-Worthy Simulation for Autonomous Cyber Operations
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.19278