CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tihanyi, Norbert, Ferrag, Mohamed Amine, Jain, Ridhi, Bisztray, Tamas, Debbah, Merouane |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2024)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2024)
How secure is AI-generated Code: A Large-Scale Comparison of Large Language Models
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2024)
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2024)
Securing Tomorrow's Smart Cities: Investigating Software Security in Internet of Vehicles and Deep Learning Technologies
von: Jain, Ridhi, et al.
Veröffentlicht: (2024)
von: Jain, Ridhi, et al.
Veröffentlicht: (2024)
The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2025)
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2025)
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
von: Dubniczky, Richard A., et al.
Veröffentlicht: (2025)
von: Dubniczky, Richard A., et al.
Veröffentlicht: (2025)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response
von: Cherif, Bilel, et al.
Veröffentlicht: (2025)
von: Cherif, Bilel, et al.
Veröffentlicht: (2025)
Revolutionizing Cyber Threat Detection with Large Language Models: A privacy-preserving BERT-based Lightweight Model for IoT/IIoT Devices
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2023)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2023)
$α^3$-SecBench: A Large-Scale Evaluation Suite of Security, Resilience, and Trust for LLM-based UAV Agents over 6G Networks
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2026)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2026)
SecureFalcon: Are We There Yet in Automated Software Vulnerability Detection with LLMs?
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2023)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2023)
Edge Learning for 6G-enabled Internet of Things: A Comprehensive Survey of Vulnerabilities, Datasets, and Defenses
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2023)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2023)
You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models
von: Dubniczky, Richard A., et al.
Veröffentlicht: (2025)
von: Dubniczky, Richard A., et al.
Veröffentlicht: (2025)
The FormAI Dataset: Generative AI in Software Security Through the Lens of Formal Verification
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2023)
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2023)
Reasoning Beyond Limits: Advances and Open Problems for LLMs
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
GenDFIR: Advancing Cyber Incident Timeline Analysis Through Retrieval Augmented Generation and Large Language Models
von: Loumachi, Fatma Yasmine, et al.
Veröffentlicht: (2024)
von: Loumachi, Fatma Yasmine, et al.
Veröffentlicht: (2024)
The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs
von: Toth, Rebeka, et al.
Veröffentlicht: (2025)
von: Toth, Rebeka, et al.
Veröffentlicht: (2025)
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge
von: Keppler, Gustav, et al.
Veröffentlicht: (2026)
von: Keppler, Gustav, et al.
Veröffentlicht: (2026)
Innovating Augmented Reality Security: Recent E2E Encryption Approaches
von: Alsop, Hamish, et al.
Veröffentlicht: (2025)
von: Alsop, Hamish, et al.
Veröffentlicht: (2025)
Reliability and Resilience of AI-Driven Critical Network Infrastructure under Cyber-Physical Threats
von: Lizos, Konstantinos A., et al.
Veröffentlicht: (2025)
von: Lizos, Konstantinos A., et al.
Veröffentlicht: (2025)
Sustaining Cyber Awareness: The Long-Term Impact of Continuous Phishing Training and Emotional Triggers
von: Toth, Rebeka, et al.
Veröffentlicht: (2025)
von: Toth, Rebeka, et al.
Veröffentlicht: (2025)
Vulnerability Detection: From Formal Verification to Large Language Models and Hybrid Approaches: A Comprehensive Overview
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2025)
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2025)
Benchmarking Knowledge-Extraction Attack and Defense on Retrieval-Augmented Generation
von: Qi, Zhisheng, et al.
Veröffentlicht: (2026)
von: Qi, Zhisheng, et al.
Veröffentlicht: (2026)
CyberLLM-FINDS 2025: Instruction-Tuned Fine-tuning of Domain-Specific LLMs with Retrieval-Augmented Generation and Graph Integration for MITRE Evaluation
von: Iyer, Vasanth, et al.
Veröffentlicht: (2026)
von: Iyer, Vasanth, et al.
Veröffentlicht: (2026)
Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents
von: Sanz-Gómez, María, et al.
Veröffentlicht: (2025)
von: Sanz-Gómez, María, et al.
Veröffentlicht: (2025)
Critical Infrastructure Protection: Generative AI, Challenges, and Opportunities
von: Yigit, Yagmur, et al.
Veröffentlicht: (2024)
von: Yigit, Yagmur, et al.
Veröffentlicht: (2024)
AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2026)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2026)
CyberAlly: Leveraging LLMs and Knowledge Graphs to Empower Cyber Defenders
von: Kim, Minjune, et al.
Veröffentlicht: (2025)
von: Kim, Minjune, et al.
Veröffentlicht: (2025)
UAVBench: An Open Benchmark Dataset for Autonomous and Agentic AI UAV Systems via LLM-Generated Flight Scenarios
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
von: Jing, Pengfei, et al.
Veröffentlicht: (2024)
von: Jing, Pengfei, et al.
Veröffentlicht: (2024)
Towards Automated Generation of Smart Grid Cyber Range for Cybersecurity Experiments and Training
von: Mashima, Daisuke, et al.
Veröffentlicht: (2024)
von: Mashima, Daisuke, et al.
Veröffentlicht: (2024)
Dynamic Intelligence Assessment: Benchmarking LLMs on the Road to AGI with a Focus on Model Confidence
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2024)
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2024)
TRACE: Timely Retrieval and Alignment for Cybersecurity Knowledge Graph Construction and Expansion
von: Xu, Zijing, et al.
Veröffentlicht: (2026)
von: Xu, Zijing, et al.
Veröffentlicht: (2026)
TechniqueRAG: Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text
von: Lekssays, Ahmed, et al.
Veröffentlicht: (2025)
von: Lekssays, Ahmed, et al.
Veröffentlicht: (2025)
CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
von: Alam, Md Tanvirul, et al.
Veröffentlicht: (2024)
von: Alam, Md Tanvirul, et al.
Veröffentlicht: (2024)
Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generation
von: Borah, Arnabh, et al.
Veröffentlicht: (2025)
von: Borah, Arnabh, et al.
Veröffentlicht: (2025)
CyberMaskQA: A Privacy-Aware Benchmark for Evaluating Large Language Models in Cybersecurity Question Answering
von: Gaddi, Matilda, et al.
Veröffentlicht: (2026)
von: Gaddi, Matilda, et al.
Veröffentlicht: (2026)
Cybersecurity and Frequent Cyber Attacks on IoT Devices in Healthcare: Issues and Solutions
von: ElSayed, Zag, et al.
Veröffentlicht: (2025)
von: ElSayed, Zag, et al.
Veröffentlicht: (2025)
A Sea of Cyber Threats: Maritime Cybersecurity from the Perspective of Mariners
von: Raymaker, Anna, et al.
Veröffentlicht: (2025)
von: Raymaker, Anna, et al.
Veröffentlicht: (2025)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2024) -
How secure is AI-generated Code: A Large-Scale Comparison of Large Language Models
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2024) -
Securing Tomorrow's Smart Cities: Investigating Software Security in Internet of Vehicles and Deep Learning Technologies
von: Jain, Ridhi, et al.
Veröffentlicht: (2024) -
The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2025) -
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
von: Dubniczky, Richard A., et al.
Veröffentlicht: (2025)