Construction and Evaluation of LLM-based agents for Semi-Autonomous penetration testing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909504135757824 |
|---|---|
| author | Kobayashi, Masaya Fuchi, Masane Zanashir, Amar Yoneda, Tomonori Takagi, Tomohiro |
| author_facet | Kobayashi, Masaya Fuchi, Masane Zanashir, Amar Yoneda, Tomonori Takagi, Tomohiro |
| contents | With the emergence of high-performance large language models (LLMs) such as GPT, Claude, and Gemini, the autonomous and semi-autonomous execution of tasks has significantly advanced across various domains. However, in highly specialized fields such as cybersecurity, full autonomy remains a challenge. This difficulty primarily stems from the limitations of LLMs in reasoning capabilities and domain-specific knowledge. We propose a system that semi-autonomously executes complex cybersecurity workflows by employing multiple LLMs modules to formulate attack strategies, generate commands, and analyze results, thereby addressing the aforementioned challenges. In our experiments using Hack The Box virtual machines, we confirmed that our system can autonomously construct attack strategies, issue appropriate commands, and automate certain processes, thereby reducing the need for manual intervention. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_15506 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Construction and Evaluation of LLM-based agents for Semi-Autonomous penetration testing Kobayashi, Masaya Fuchi, Masane Zanashir, Amar Yoneda, Tomonori Takagi, Tomohiro Cryptography and Security With the emergence of high-performance large language models (LLMs) such as GPT, Claude, and Gemini, the autonomous and semi-autonomous execution of tasks has significantly advanced across various domains. However, in highly specialized fields such as cybersecurity, full autonomy remains a challenge. This difficulty primarily stems from the limitations of LLMs in reasoning capabilities and domain-specific knowledge. We propose a system that semi-autonomously executes complex cybersecurity workflows by employing multiple LLMs modules to formulate attack strategies, generate commands, and analyze results, thereby addressing the aforementioned challenges. In our experiments using Hack The Box virtual machines, we confirmed that our system can autonomously construct attack strategies, issue appropriate commands, and automate certain processes, thereby reducing the need for manual intervention. |
| title | Construction and Evaluation of LLM-based agents for Semi-Autonomous penetration testing |
| topic | Cryptography and Security |
| url | https://arxiv.org/abs/2502.15506 |