Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hariharan, Suhas, Majid, Zainab Ali, Veuthey, Jaime Raldua, Haimes, Jacob
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909388279644160
author Hariharan, Suhas
Majid, Zainab Ali
Veuthey, Jaime Raldua
Haimes, Jacob
author_facet Hariharan, Suhas
Majid, Zainab Ali
Veuthey, Jaime Raldua
Haimes, Jacob
contents A key development in the cybersecurity evaluations space is the work carried out by Meta, through their CyberSecEval approach. While this work is undoubtedly a useful contribution to a nascent field, there are notable features that limit its utility. Key drawbacks focus on the insecure code detection part of Meta's methodology. We explore these limitations, and use our exploration as a test case for LLM-assisted benchmark analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08813
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
Hariharan, Suhas
Majid, Zainab Ali
Veuthey, Jaime Raldua
Haimes, Jacob
Artificial Intelligence
A key development in the cybersecurity evaluations space is the work carried out by Meta, through their CyberSecEval approach. While this work is undoubtedly a useful contribution to a nascent field, there are notable features that limit its utility. Key drawbacks focus on the insecure code detection part of Meta's methodology. We explore these limitations, and use our exploration as a test case for LLM-assisted benchmark analysis.
title Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
topic Artificial Intelligence
url https://arxiv.org/abs/2411.08813