CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Keppler, Gustav, Elbez, Ghada, Hagenmeyer, Veit
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908986217857024
author Keppler, Gustav
Elbez, Ghada
Hagenmeyer, Veit
author_facet Keppler, Gustav
Elbez, Ghada
Hagenmeyer, Veit
contents The rapid evolution and use of Large Language Models (LLMs) in professional workflows require an evaluation of their domain-specific knowledge against industry standards. We introduceCyberCertBench, a new suite of Multiple Choice Question Answering (MCQA) benchmarks derived from industry recognized certifications. CyberCertBench evaluates LLM domain knowledgeagainst the professional standards of Information Technology cybersecurity and more specializedareas such as Operational Technology and related cybersecurity standards. Concurrently, we propose and validate a novel Proposer-Verifier framework, a methodology to generate interpretable,natural language explanations for model performance. Our evaluation shows that frontier modelsachieve human expert level in general networking and IT security knowledge. However, theiraccuracy declines in questions that require vendor-specific nuances or knowledge in formalstandards, like, e.g., IEC 62443. Analysis of model scaling trend and release date demonstratesremarkable gains in parameter efficiency, while recent larger models show diminishing returns.Code and evaluation scripts are available at: https://github.com/GKeppler/CyberCertBench.
format Preprint
id arxiv_https___arxiv_org_abs_2604_20389
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge
Keppler, Gustav
Elbez, Ghada
Hagenmeyer, Veit
Cryptography and Security
Artificial Intelligence
The rapid evolution and use of Large Language Models (LLMs) in professional workflows require an evaluation of their domain-specific knowledge against industry standards. We introduceCyberCertBench, a new suite of Multiple Choice Question Answering (MCQA) benchmarks derived from industry recognized certifications. CyberCertBench evaluates LLM domain knowledgeagainst the professional standards of Information Technology cybersecurity and more specializedareas such as Operational Technology and related cybersecurity standards. Concurrently, we propose and validate a novel Proposer-Verifier framework, a methodology to generate interpretable,natural language explanations for model performance. Our evaluation shows that frontier modelsachieve human expert level in general networking and IT security knowledge. However, theiraccuracy declines in questions that require vendor-specific nuances or knowledge in formalstandards, like, e.g., IEC 62443. Analysis of model scaling trend and release date demonstratesremarkable gains in parameter efficiency, while recent larger models show diminishing returns.Code and evaluation scripts are available at: https://github.com/GKeppler/CyberCertBench.
title CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2604.20389