An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shao, Minghao, Chen, Boyuan, Jancheska, Sofija, Dolan-Gavitt, Brendan, Garg, Siddharth, Karri, Ramesh, Shafique, Muhammad
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917593110020096
author Shao, Minghao
Chen, Boyuan
Jancheska, Sofija
Dolan-Gavitt, Brendan
Garg, Siddharth
Karri, Ramesh
Shafique, Muhammad
author_facet Shao, Minghao
Chen, Boyuan
Jancheska, Sofija
Dolan-Gavitt, Brendan
Garg, Siddharth
Karri, Ramesh
Shafique, Muhammad
contents Capture The Flag (CTF) challenges are puzzles related to computer security scenarios. With the advent of large language models (LLMs), more and more CTF participants are using LLMs to understand and solve the challenges. However, so far no work has evaluated the effectiveness of LLMs in solving CTF challenges with a fully automated workflow. We develop two CTF-solving workflows, human-in-the-loop (HITL) and fully-automated, to examine the LLMs' ability to solve a selected set of CTF challenges, prompted with information about the question. We collect human contestants' results on the same set of questions, and find that LLMs achieve higher success rate than an average human participant. This work provides a comprehensive evaluation of the capability of LLMs in solving real world CTF challenges, from real competition to fully automated workflow. Our results provide references for applying LLMs in cybersecurity education and pave the way for systematic evaluation of offensive cybersecurity capabilities in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11814
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
Shao, Minghao
Chen, Boyuan
Jancheska, Sofija
Dolan-Gavitt, Brendan
Garg, Siddharth
Karri, Ramesh
Shafique, Muhammad
Cryptography and Security
Capture The Flag (CTF) challenges are puzzles related to computer security scenarios. With the advent of large language models (LLMs), more and more CTF participants are using LLMs to understand and solve the challenges. However, so far no work has evaluated the effectiveness of LLMs in solving CTF challenges with a fully automated workflow. We develop two CTF-solving workflows, human-in-the-loop (HITL) and fully-automated, to examine the LLMs' ability to solve a selected set of CTF challenges, prompted with information about the question. We collect human contestants' results on the same set of questions, and find that LLMs achieve higher success rate than an average human participant. This work provides a comprehensive evaluation of the capability of LLMs in solving real world CTF challenges, from real competition to fully automated workflow. Our results provide references for applying LLMs in cybersecurity education and pave the way for systematic evaluation of offensive cybersecurity capabilities in LLMs.
title An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
topic Cryptography and Security
url https://arxiv.org/abs/2402.11814