CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | Keppler, Gustav, Elbez, Ghada, Hagenmeyer, Veit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Global Analysis of Cyber Threats to the Energy Sector: "Currents of Conflict" from a Geopolitical Perspective
by: Sánchez, Gustavo, et al.
Published: (2025)
by: Sánchez, Gustavo, et al.
Published: (2025)
CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments
by: Keppler, Gustav, et al.
Published: (2026)
by: Keppler, Gustav, et al.
Published: (2026)
RTS-ABAC: Real-Time Server-Aided Attribute-Based Authorization & Access Control for Substation Automation Systems
by: Gstür, Moritz, et al.
Published: (2026)
by: Gstür, Moritz, et al.
Published: (2026)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
FullCert: Deterministic End-to-End Certification for Training and Inference of Neural Networks
by: Lorenz, Tobias, et al.
Published: (2024)
by: Lorenz, Tobias, et al.
Published: (2024)
AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
by: Jing, Pengfei, et al.
Published: (2024)
by: Jing, Pengfei, et al.
Published: (2024)
CrossCert: A Cross-Checking Detection Approach to Patch Robustness Certification for Deep Learning Models
by: Zhou, Qilin, et al.
Published: (2024)
by: Zhou, Qilin, et al.
Published: (2024)
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
by: Fan, Yihe, et al.
Published: (2026)
by: Fan, Yihe, et al.
Published: (2026)
CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2024)
by: Alam, Md Tanvirul, et al.
Published: (2024)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
by: Lee, Seunghyun, et al.
Published: (2026)
by: Lee, Seunghyun, et al.
Published: (2026)
Artificial Intelligence in Cybersecurity: Building Resilient Cyber Diplomacy Frameworks
by: Stoltz, Michael
Published: (2024)
by: Stoltz, Michael
Published: (2024)
When LLMs Meet Cybersecurity: A Systematic Literature Review
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
by: Lim, Taein, et al.
Published: (2026)
by: Lim, Taein, et al.
Published: (2026)
AgenticCyber: A GenAI-Powered Multi-Agent System for Multimodal Threat Detection and Adaptive Response in Cybersecurity
by: Roy, Shovan
Published: (2025)
by: Roy, Shovan
Published: (2025)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
by: Wu, Yiran, et al.
Published: (2025)
by: Wu, Yiran, et al.
Published: (2025)
Contextualized AI for Cyber Defense: An Automated Survey using LLMs
by: Haryanto, Christoforus Yoga, et al.
Published: (2024)
by: Haryanto, Christoforus Yoga, et al.
Published: (2024)
CyberCScope: Mining Skewed Tensor Streams and Online Anomaly Detection in Cybersecurity Systems
by: Nakamura, Kota, et al.
Published: (2025)
by: Nakamura, Kota, et al.
Published: (2025)
AEGIS: White-Box Attack Path Generation using LLMs and Training Effectiveness Evaluation for Large-Scale Cyber Defence Exercises
by: Tung, Ivan K., et al.
Published: (2026)
by: Tung, Ivan K., et al.
Published: (2026)
CurricuLLM: Designing Personalized and Workforce-Aligned Cybersecurity Curricula Using Fine-Tuned LLMs
by: Nijdam, Arthur, et al.
Published: (2026)
by: Nijdam, Arthur, et al.
Published: (2026)
CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
by: Deason, Lauren, et al.
Published: (2025)
by: Deason, Lauren, et al.
Published: (2025)
CTIArena: Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence
by: Cheng, Yutong, et al.
Published: (2025)
by: Cheng, Yutong, et al.
Published: (2025)
Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks
by: Dahiya, Vivek, et al.
Published: (2026)
by: Dahiya, Vivek, et al.
Published: (2026)
Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data
by: ElZemity, Adel, et al.
Published: (2025)
by: ElZemity, Adel, et al.
Published: (2025)
Integrative Approaches in Cybersecurity and AI
by: Omar, Marwan
Published: (2024)
by: Omar, Marwan
Published: (2024)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
by: Kouremetis, Michael, et al.
Published: (2025)
by: Kouremetis, Michael, et al.
Published: (2025)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
by: Ma, Haokai, et al.
Published: (2025)
by: Ma, Haokai, et al.
Published: (2025)
Leveraging Large Language Models for Cybersecurity Risk Assessment -- A Case from Forestry Cyber-Physical Systems
by: Gultekin, Fikret Mert, et al.
Published: (2025)
by: Gultekin, Fikret Mert, et al.
Published: (2025)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
by: Chen, Qirui, et al.
Published: (2026)
by: Chen, Qirui, et al.
Published: (2026)
Quantitative Certification of Agentic Tool Selection
by: Yeon, Jehyeok, et al.
Published: (2025)
by: Yeon, Jehyeok, et al.
Published: (2025)
A Survey on Offensive AI Within Cybersecurity
by: Girhepuje, Sahil, et al.
Published: (2024)
by: Girhepuje, Sahil, et al.
Published: (2024)
A Survey of Large Language Models in Cybersecurity
by: da Silva, Gabriel de Jesus Coelho, et al.
Published: (2024)
by: da Silva, Gabriel de Jesus Coelho, et al.
Published: (2024)
Reinforcement Learning for Automated Cybersecurity Penetration Testing
by: López-Montero, Daniel, et al.
Published: (2025)
by: López-Montero, Daniel, et al.
Published: (2025)
Dynamic Risk Assessments for Offensive Cybersecurity Agents
by: Wei, Boyi, et al.
Published: (2025)
by: Wei, Boyi, et al.
Published: (2025)
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
by: Ristea, Dan, et al.
Published: (2024)
by: Ristea, Dan, et al.
Published: (2024)
Beyond RAG for Cyber Threat Intelligence: A Systematic Evaluation of Graph-Based and Agentic Retrieval
by: Hamzic, Dzenan, et al.
Published: (2026)
by: Hamzic, Dzenan, et al.
Published: (2026)
The Path To Autonomous Cyber Defense
by: Oesch, Sean, et al.
Published: (2024)
by: Oesch, Sean, et al.
Published: (2024)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
by: Black, Sid, et al.
Published: (2025)
by: Black, Sid, et al.
Published: (2025)
Similar Items
-
A Global Analysis of Cyber Threats to the Energy Sector: "Currents of Conflict" from a Geopolitical Perspective
by: Sánchez, Gustavo, et al.
Published: (2025) -
CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments
by: Keppler, Gustav, et al.
Published: (2026) -
RTS-ABAC: Real-Time Server-Aided Attribute-Based Authorization & Access Control for Substation Automation Systems
by: Gstür, Moritz, et al.
Published: (2026) -
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024) -
FullCert: Deterministic End-to-End Certification for Training and Inference of Neural Networks
by: Lorenz, Tobias, et al.
Published: (2024)