CIPHER: Cybersecurity Intelligent Penetration-testing Helper for Ethical Researcher

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pratama, Derry, Suryanto, Naufal, Adiputra, Andro Aprila, Le, Thi-Thu-Huong, Kadiptya, Ahmada Yusril, Iqbal, Muhammad, Kim, Howon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929578499375104
author Pratama, Derry
Suryanto, Naufal
Adiputra, Andro Aprila
Le, Thi-Thu-Huong
Kadiptya, Ahmada Yusril
Iqbal, Muhammad
Kim, Howon
author_facet Pratama, Derry
Suryanto, Naufal
Adiputra, Andro Aprila
Le, Thi-Thu-Huong
Kadiptya, Ahmada Yusril
Iqbal, Muhammad
Kim, Howon
contents Penetration testing, a critical component of cybersecurity, typically requires extensive time and effort to find vulnerabilities. Beginners in this field often benefit from collaborative approaches with the community or experts. To address this, we develop CIPHER (Cybersecurity Intelligent Penetration-testing Helper for Ethical Researchers), a large language model specifically trained to assist in penetration testing tasks. We trained CIPHER using over 300 high-quality write-ups of vulnerable machines, hacking techniques, and documentation of open-source penetration testing tools. Additionally, we introduced the Findings, Action, Reasoning, and Results (FARR) Flow augmentation, a novel method to augment penetration testing write-ups to establish a fully automated pentesting simulation benchmark tailored for large language models. This approach fills a significant gap in traditional cybersecurity Q\&A benchmarks and provides a realistic and rigorous standard for evaluating AI's technical knowledge, reasoning capabilities, and practical utility in dynamic penetration testing scenarios. In our assessments, CIPHER achieved the best overall performance in providing accurate suggestion responses compared to other open-source penetration testing models of similar size and even larger state-of-the-art models like Llama 3 70B and Qwen1.5 72B Chat, particularly on insane difficulty machine setups. This demonstrates that the current capabilities of general LLMs are insufficient for effectively guiding users through the penetration testing process. We also discuss the potential for improvement through scaling and the development of better benchmarks using FARR Flow augmentation results. Our benchmark will be released publicly at https://github.com/ibndias/CIPHER.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11650
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CIPHER: Cybersecurity Intelligent Penetration-testing Helper for Ethical Researcher
Pratama, Derry
Suryanto, Naufal
Adiputra, Andro Aprila
Le, Thi-Thu-Huong
Kadiptya, Ahmada Yusril
Iqbal, Muhammad
Kim, Howon
Cryptography and Security
Artificial Intelligence
Penetration testing, a critical component of cybersecurity, typically requires extensive time and effort to find vulnerabilities. Beginners in this field often benefit from collaborative approaches with the community or experts. To address this, we develop CIPHER (Cybersecurity Intelligent Penetration-testing Helper for Ethical Researchers), a large language model specifically trained to assist in penetration testing tasks. We trained CIPHER using over 300 high-quality write-ups of vulnerable machines, hacking techniques, and documentation of open-source penetration testing tools. Additionally, we introduced the Findings, Action, Reasoning, and Results (FARR) Flow augmentation, a novel method to augment penetration testing write-ups to establish a fully automated pentesting simulation benchmark tailored for large language models. This approach fills a significant gap in traditional cybersecurity Q\&A benchmarks and provides a realistic and rigorous standard for evaluating AI's technical knowledge, reasoning capabilities, and practical utility in dynamic penetration testing scenarios. In our assessments, CIPHER achieved the best overall performance in providing accurate suggestion responses compared to other open-source penetration testing models of similar size and even larger state-of-the-art models like Llama 3 70B and Qwen1.5 72B Chat, particularly on insane difficulty machine setups. This demonstrates that the current capabilities of general LLMs are insufficient for effectively guiding users through the penetration testing process. We also discuss the potential for improvement through scaling and the development of better benchmarks using FARR Flow augmentation results. Our benchmark will be released publicly at https://github.com/ibndias/CIPHER.
title CIPHER: Cybersecurity Intelligent Penetration-testing Helper for Ethical Researcher
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2408.11650