AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Alam, Md Tanvirul, Bhusal, Dipkamal, Ahmad, Salman, Rastogi, Nidhi, Worth, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2024)
by: Alam, Md Tanvirul, et al.
Published: (2024)
Actionable Cyber Threat Intelligence using Knowledge Graphs and Large Language Models
by: Fieblinger, Romy, et al.
Published: (2024)
by: Fieblinger, Romy, et al.
Published: (2024)
R+R: Revisiting Static Feature-Based Android Malware Detection using Machine Learning
by: Alam, Md Tanvirul, et al.
Published: (2024)
by: Alam, Md Tanvirul, et al.
Published: (2024)
PASA: Attack Agnostic Unsupervised Adversarial Detection using Prediction & Attribution Sensitivity Analysis
by: Bhusal, Dipkamal, et al.
Published: (2024)
by: Bhusal, Dipkamal, et al.
Published: (2024)
Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generation
by: Borah, Arnabh, et al.
Published: (2025)
by: Borah, Arnabh, et al.
Published: (2025)
Adversarial Patterns: Building Robust Android Malware Classifiers
by: Bhusal, Dipkamal, et al.
Published: (2022)
by: Bhusal, Dipkamal, et al.
Published: (2022)
SECURE: Benchmarking Large Language Models for Cybersecurity
by: Bhusal, Dipkamal, et al.
Published: (2024)
by: Bhusal, Dipkamal, et al.
Published: (2024)
ADAPT: A Pseudo-labeling Approach to Combat Concept Drift in Malware Detection
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
by: Deason, Lauren, et al.
Published: (2025)
by: Deason, Lauren, et al.
Published: (2025)
Cyber Threat Intelligence for Artificial Intelligence Systems
by: Krawczyk, Natalia, et al.
Published: (2026)
by: Krawczyk, Natalia, et al.
Published: (2026)
CTIArena: Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence
by: Cheng, Yutong, et al.
Published: (2025)
by: Cheng, Yutong, et al.
Published: (2025)
PROVEX: Enhancing SOC Analyst Trust with Explainable Provenance-Based IDS
by: Dhanuka, Devang, et al.
Published: (2025)
by: Dhanuka, Devang, et al.
Published: (2025)
CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge
by: Keppler, Gustav, et al.
Published: (2026)
by: Keppler, Gustav, et al.
Published: (2026)
Training for Trustworthy Saliency Maps: Adversarial Training Meets Feature-Map Smoothing
by: Bhusal, Dipkamal, et al.
Published: (2026)
by: Bhusal, Dipkamal, et al.
Published: (2026)
Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence
by: Meng, Yuqiao, et al.
Published: (2025)
by: Meng, Yuqiao, et al.
Published: (2025)
Beyond RAG for Cyber Threat Intelligence: A Systematic Evaluation of Graph-Based and Agentic Retrieval
by: Hamzic, Dzenan, et al.
Published: (2026)
by: Hamzic, Dzenan, et al.
Published: (2026)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
by: Wu, Yiran, et al.
Published: (2025)
by: Wu, Yiran, et al.
Published: (2025)
XAI-CF -- Examining the Role of Explainable Artificial Intelligence in Cyber Forensics
by: Alam, Shahid, et al.
Published: (2024)
by: Alam, Shahid, et al.
Published: (2024)
AI-Driven Cyber Threat Intelligence Automation
by: Shah, Shrit, et al.
Published: (2024)
by: Shah, Shrit, et al.
Published: (2024)
Large Language Models Are Unreliable for Cyber Threat Intelligence
by: Mezzi, Emanuele, et al.
Published: (2025)
by: Mezzi, Emanuele, et al.
Published: (2025)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
by: Lim, Taein, et al.
Published: (2026)
by: Lim, Taein, et al.
Published: (2026)
FALCON: Autonomous Cyber Threat Intelligence Mining with LLMs for IDS Rule Generation
by: Mitra, Shaswata, et al.
Published: (2025)
by: Mitra, Shaswata, et al.
Published: (2025)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
by: Chona, Alankrit, et al.
Published: (2026)
by: Chona, Alankrit, et al.
Published: (2026)
TrojanWhisper: Evaluating Pre-trained LLMs to Detect and Localize Hardware Trojans
by: Faruque, Md Omar, et al.
Published: (2024)
by: Faruque, Md Omar, et al.
Published: (2024)
Minerva: Reinforcement Learning with Verifiable Rewards for Cyber Threat Intelligence LLMs
by: Alam, Md Tanvirul, et al.
Published: (2026)
by: Alam, Md Tanvirul, et al.
Published: (2026)
CyberSentinel: An Emergent Threat Detection System for AI Security
by: Tallam, Krti
Published: (2025)
by: Tallam, Krti
Published: (2025)
Countering Autonomous Cyber Threats
by: Heckel, Kade M., et al.
Published: (2024)
by: Heckel, Kade M., et al.
Published: (2024)
POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment
by: Tang, Luoxi, et al.
Published: (2025)
by: Tang, Luoxi, et al.
Published: (2025)
BARTPredict: Empowering IoT Security with LLM-Driven Cyber Threat Prediction
by: Diaf, Alaeddine, et al.
Published: (2025)
by: Diaf, Alaeddine, et al.
Published: (2025)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
by: Ma, Haokai, et al.
Published: (2025)
by: Ma, Haokai, et al.
Published: (2025)
TITAN: Graph-Executable Reasoning for Cyber Threat Intelligence
by: Simoni, Marco, et al.
Published: (2025)
by: Simoni, Marco, et al.
Published: (2025)
The Use of Large Language Models (LLM) for Cyber Threat Intelligence (CTI) in Cybercrime Forums
by: Clairoux-Trepanier, Vanessa, et al.
Published: (2024)
by: Clairoux-Trepanier, Vanessa, et al.
Published: (2024)
Technique Inference Engine: A Recommender Model to Support Cyber Threat Hunting
by: Turner, Matthew J., et al.
Published: (2025)
by: Turner, Matthew J., et al.
Published: (2025)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
by: Chen, Qirui, et al.
Published: (2026)
by: Chen, Qirui, et al.
Published: (2026)
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
by: Jing, Pengfei, et al.
Published: (2024)
by: Jing, Pengfei, et al.
Published: (2024)
Mind the Gap: Missing Cyber Threat Coverage in NIDS Datasets for the Energy Sector
by: Tory, Adrita Rahman, et al.
Published: (2025)
by: Tory, Adrita Rahman, et al.
Published: (2025)
CTINexus: Automatic Cyber Threat Intelligence Knowledge Graph Construction Using Large Language Models
by: Cheng, Yutong, et al.
Published: (2024)
by: Cheng, Yutong, et al.
Published: (2024)
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
by: Ristea, Dan, et al.
Published: (2024)
by: Ristea, Dan, et al.
Published: (2024)
Towards Explainable and Lightweight AI for Real-Time Cyber Threat Hunting in Edge Networks
by: Rahmati, Milad
Published: (2025)
by: Rahmati, Milad
Published: (2025)
Similar Items
-
CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2024) -
Actionable Cyber Threat Intelligence using Knowledge Graphs and Large Language Models
by: Fieblinger, Romy, et al.
Published: (2024) -
R+R: Revisiting Static Feature-Based Android Malware Detection using Machine Learning
by: Alam, Md Tanvirul, et al.
Published: (2024) -
PASA: Attack Agnostic Unsupervised Adversarial Detection using Prediction & Attribution Sensitivity Analysis
by: Bhusal, Dipkamal, et al.
Published: (2024) -
Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generation
by: Borah, Arnabh, et al.
Published: (2025)