Evaluating LLM Generated Detection Rules in Cybersecurity
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916959047647232 |
|---|---|
| author | Bertiger, Anna Filar, Bobby Luthra, Aryan Meschiari, Stefano Mitchell, Aiden Scholten, Sam Sharath, Vivek |
| author_facet | Bertiger, Anna Filar, Bobby Luthra, Aryan Meschiari, Stefano Mitchell, Aiden Scholten, Sam Sharath, Vivek |
| contents | LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners. Here, we present an open-source evaluation framework and benchmark metrics for evaluating LLM-generated cybersecurity rules. The benchmark employs a holdout set-based methodology to measure the effectiveness of LLM-generated security rules in comparison to a human-generated corpus of rules. It provides three key metrics inspired by the way experts evaluate security rules, offering a realistic, multifaceted evaluation of the effectiveness of an LLM-based security rule generator. This methodology is illustrated using rules from Sublime Security's detection team and those written by Sublime Security's Automated Detection Engineer (ADE), with a thorough analysis of ADE's skills presented in the results section. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_16749 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Evaluating LLM Generated Detection Rules in Cybersecurity Bertiger, Anna Filar, Bobby Luthra, Aryan Meschiari, Stefano Mitchell, Aiden Scholten, Sam Sharath, Vivek Cryptography and Security LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners. Here, we present an open-source evaluation framework and benchmark metrics for evaluating LLM-generated cybersecurity rules. The benchmark employs a holdout set-based methodology to measure the effectiveness of LLM-generated security rules in comparison to a human-generated corpus of rules. It provides three key metrics inspired by the way experts evaluate security rules, offering a realistic, multifaceted evaluation of the effectiveness of an LLM-based security rule generator. This methodology is illustrated using rules from Sublime Security's detection team and those written by Sublime Security's Automated Detection Engineer (ADE), with a thorough analysis of ADE's skills presented in the results section. |
| title | Evaluating LLM Generated Detection Rules in Cybersecurity |
| topic | Cryptography and Security |
| url | https://arxiv.org/abs/2509.16749 |