Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909388279644160 |
|---|---|
| author | Hariharan, Suhas Majid, Zainab Ali Veuthey, Jaime Raldua Haimes, Jacob |
| author_facet | Hariharan, Suhas Majid, Zainab Ali Veuthey, Jaime Raldua Haimes, Jacob |
| contents | A key development in the cybersecurity evaluations space is the work carried out by Meta, through their CyberSecEval approach. While this work is undoubtedly a useful contribution to a nascent field, there are notable features that limit its utility. Key drawbacks focus on the insecure code detection part of Meta's methodology. We explore these limitations, and use our exploration as a test case for LLM-assisted benchmark analysis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_08813 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique Hariharan, Suhas Majid, Zainab Ali Veuthey, Jaime Raldua Haimes, Jacob Artificial Intelligence A key development in the cybersecurity evaluations space is the work carried out by Meta, through their CyberSecEval approach. While this work is undoubtedly a useful contribution to a nascent field, there are notable features that limit its utility. Key drawbacks focus on the insecure code detection part of Meta's methodology. We explore these limitations, and use our exploration as a test case for LLM-assisted benchmark analysis. |
| title | Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2411.08813 |