Evaluating LLM Generated Detection Rules in Cybersecurity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bertiger, Anna, Filar, Bobby, Luthra, Aryan, Meschiari, Stefano, Mitchell, Aiden, Scholten, Sam, Sharath, Vivek
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916959047647232
author Bertiger, Anna
Filar, Bobby
Luthra, Aryan
Meschiari, Stefano
Mitchell, Aiden
Scholten, Sam
Sharath, Vivek
author_facet Bertiger, Anna
Filar, Bobby
Luthra, Aryan
Meschiari, Stefano
Mitchell, Aiden
Scholten, Sam
Sharath, Vivek
contents LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners. Here, we present an open-source evaluation framework and benchmark metrics for evaluating LLM-generated cybersecurity rules. The benchmark employs a holdout set-based methodology to measure the effectiveness of LLM-generated security rules in comparison to a human-generated corpus of rules. It provides three key metrics inspired by the way experts evaluate security rules, offering a realistic, multifaceted evaluation of the effectiveness of an LLM-based security rule generator. This methodology is illustrated using rules from Sublime Security's detection team and those written by Sublime Security's Automated Detection Engineer (ADE), with a thorough analysis of ADE's skills presented in the results section.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating LLM Generated Detection Rules in Cybersecurity
Bertiger, Anna
Filar, Bobby
Luthra, Aryan
Meschiari, Stefano
Mitchell, Aiden
Scholten, Sam
Sharath, Vivek
Cryptography and Security
LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners. Here, we present an open-source evaluation framework and benchmark metrics for evaluating LLM-generated cybersecurity rules. The benchmark employs a holdout set-based methodology to measure the effectiveness of LLM-generated security rules in comparison to a human-generated corpus of rules. It provides three key metrics inspired by the way experts evaluate security rules, offering a realistic, multifaceted evaluation of the effectiveness of an LLM-based security rule generator. This methodology is illustrated using rules from Sublime Security's detection team and those written by Sublime Security's Automated Detection Engineer (ADE), with a thorough analysis of ADE's skills presented in the results section.
title Evaluating LLM Generated Detection Rules in Cybersecurity
topic Cryptography and Security
url https://arxiv.org/abs/2509.16749