AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xi, Haoran, Shao, Minghao, Milner, Kimberly, Putrevu, Venkata Sai Charan, Rani, Nanda, Udeshi, Meet, Krishnamurthy, Prashanth, Dolan-Gavitt, Brendan, Garg, Siddharth, Shukla, Sandeep Kumar, Khorrami, Farshad, Hillel-Tuch, Alon, Shafique, Muhammad, Karri, Ramesh
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911557478252544
author Xi, Haoran
Shao, Minghao
Milner, Kimberly
Putrevu, Venkata Sai Charan
Rani, Nanda
Udeshi, Meet
Krishnamurthy, Prashanth
Dolan-Gavitt, Brendan
Garg, Siddharth
Shukla, Sandeep Kumar
Khorrami, Farshad
Hillel-Tuch, Alon
Shafique, Muhammad
Karri, Ramesh
author_facet Xi, Haoran
Shao, Minghao
Milner, Kimberly
Putrevu, Venkata Sai Charan
Rani, Nanda
Udeshi, Meet
Krishnamurthy, Prashanth
Dolan-Gavitt, Brendan
Garg, Siddharth
Shukla, Sandeep Kumar
Khorrami, Farshad
Hillel-Tuch, Alon
Shafique, Muhammad
Karri, Ramesh
contents Large language models are rapidly changing how learners acquire and demonstrate cybersecurity skills. However, when human--AI collaboration is allowed, educators still lack validated competition designs and evaluation practices that remain fair and evidence-based. This paper presents a cross-regional study of LLM-centered Capture-the-Flag competitions built on the Cyber Security Awareness Week competition system. To understand how autonomy levels and participants' knowledge backgrounds influence problem-solving performance and learning-related behaviors, we formalize three autonomy levels: human-in-the-loop, autonomous agent frameworks, and hybrid. To enable verification, we require traceable submissions including conversation logs, agent trajectories, and agent code. We analyze multi-region competition data covering an in-class track, a standard track, and a year-long expert track, each targeting participants with different knowledge backgrounds. Using data from the 2025 competition, we compare solve performance across autonomy levels and challenge categories, and observe that autonomous agent frameworks and hybrid achieve higher completion rates on challenges requiring iterative testing and tool interactions. In the in-class track, we classify participants' agent designs and find a preference for lightweight, tool-augmented prompting and reflection-based retries over complex multi-agent architectures. Our results offer actionable guidance for designing LLM-assisted cybersecurity competitions as learning technologies, including autonomy-specific scoring criteria, evidence requirements that support solution verification, and track structures that improve accessibility while preserving reliable evaluation and engagement.
format Preprint
id arxiv_https___arxiv_org_abs_2603_21551
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes
Xi, Haoran
Shao, Minghao
Milner, Kimberly
Putrevu, Venkata Sai Charan
Rani, Nanda
Udeshi, Meet
Krishnamurthy, Prashanth
Dolan-Gavitt, Brendan
Garg, Siddharth
Shukla, Sandeep Kumar
Khorrami, Farshad
Hillel-Tuch, Alon
Shafique, Muhammad
Karri, Ramesh
Software Engineering
Large language models are rapidly changing how learners acquire and demonstrate cybersecurity skills. However, when human--AI collaboration is allowed, educators still lack validated competition designs and evaluation practices that remain fair and evidence-based. This paper presents a cross-regional study of LLM-centered Capture-the-Flag competitions built on the Cyber Security Awareness Week competition system. To understand how autonomy levels and participants' knowledge backgrounds influence problem-solving performance and learning-related behaviors, we formalize three autonomy levels: human-in-the-loop, autonomous agent frameworks, and hybrid. To enable verification, we require traceable submissions including conversation logs, agent trajectories, and agent code. We analyze multi-region competition data covering an in-class track, a standard track, and a year-long expert track, each targeting participants with different knowledge backgrounds. Using data from the 2025 competition, we compare solve performance across autonomy levels and challenge categories, and observe that autonomous agent frameworks and hybrid achieve higher completion rates on challenges requiring iterative testing and tool interactions. In the in-class track, we classify participants' agent designs and find a preference for lightweight, tool-augmented prompting and reflection-based retries over complex multi-agent architectures. Our results offer actionable guidance for designing LLM-assisted cybersecurity competitions as learning technologies, including autonomy-specific scoring criteria, evidence requirements that support solution verification, and track structures that improve accessibility while preserving reliable evaluation and engagement.
title AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes
topic Software Engineering
url https://arxiv.org/abs/2603.21551