NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shao, Minghao, Jancheska, Sofija, Udeshi, Meet, Dolan-Gavitt, Brendan, Xi, Haoran, Milner, Kimberly, Chen, Boyuan, Yin, Max, Garg, Siddharth, Krishnamurthy, Prashanth, Khorrami, Farshad, Karri, Ramesh, Shafique, Muhammad
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915156225687552
author Shao, Minghao
Jancheska, Sofija
Udeshi, Meet
Dolan-Gavitt, Brendan
Xi, Haoran
Milner, Kimberly
Chen, Boyuan
Yin, Max
Garg, Siddharth
Krishnamurthy, Prashanth
Khorrami, Farshad
Karri, Ramesh
Shafique, Muhammad
author_facet Shao, Minghao
Jancheska, Sofija
Udeshi, Meet
Dolan-Gavitt, Brendan
Xi, Haoran
Milner, Kimberly
Chen, Boyuan
Yin, Max
Garg, Siddharth
Krishnamurthy, Prashanth
Khorrami, Farshad
Karri, Ramesh
Shafique, Muhammad
contents Large Language Models (LLMs) are being deployed across various domains today. However, their capacity to solve Capture the Flag (CTF) challenges in cybersecurity has not been thoroughly evaluated. To address this, we develop a novel method to assess LLMs in solving CTF challenges by creating a scalable, open-source benchmark database specifically designed for these applications. This database includes metadata for LLM testing and adaptive learning, compiling a diverse range of CTF challenges from popular competitions. Utilizing the advanced function calling capabilities of LLMs, we build a fully automated system with an enhanced workflow and support for external tool calls. Our benchmark dataset and automated framework allow us to evaluate the performance of five LLMs, encompassing both black-box and open-source models. This work lays the foundation for future research into improving the efficiency of LLMs in interactive cybersecurity tasks and automated task planning. By providing a specialized benchmark, our project offers an ideal platform for developing, testing, and refining LLM-based approaches to vulnerability detection and resolution. Evaluating LLMs on these challenges and comparing with human performance yields insights into their potential for AI-driven cybersecurity solutions to perform real-world threat management. We make our benchmark dataset open source to public https://github.com/NYU-LLM-CTF/NYU_CTF_Bench along with our playground automated framework https://github.com/NYU-LLM-CTF/llm_ctf_automation.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05590
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
Shao, Minghao
Jancheska, Sofija
Udeshi, Meet
Dolan-Gavitt, Brendan
Xi, Haoran
Milner, Kimberly
Chen, Boyuan
Yin, Max
Garg, Siddharth
Krishnamurthy, Prashanth
Khorrami, Farshad
Karri, Ramesh
Shafique, Muhammad
Cryptography and Security
Artificial Intelligence
Computers and Society
Machine Learning
Large Language Models (LLMs) are being deployed across various domains today. However, their capacity to solve Capture the Flag (CTF) challenges in cybersecurity has not been thoroughly evaluated. To address this, we develop a novel method to assess LLMs in solving CTF challenges by creating a scalable, open-source benchmark database specifically designed for these applications. This database includes metadata for LLM testing and adaptive learning, compiling a diverse range of CTF challenges from popular competitions. Utilizing the advanced function calling capabilities of LLMs, we build a fully automated system with an enhanced workflow and support for external tool calls. Our benchmark dataset and automated framework allow us to evaluate the performance of five LLMs, encompassing both black-box and open-source models. This work lays the foundation for future research into improving the efficiency of LLMs in interactive cybersecurity tasks and automated task planning. By providing a specialized benchmark, our project offers an ideal platform for developing, testing, and refining LLM-based approaches to vulnerability detection and resolution. Evaluating LLMs on these challenges and comparing with human performance yields insights into their potential for AI-driven cybersecurity solutions to perform real-world threat management. We make our benchmark dataset open source to public https://github.com/NYU-LLM-CTF/NYU_CTF_Bench along with our playground automated framework https://github.com/NYU-LLM-CTF/llm_ctf_automation.
title NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
topic Cryptography and Security
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2406.05590