Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dahiya, Vivek, Nehra, Sunny, Dholariya, Vipul, Shangari, Bhavik, Khatri, Chandra |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
von: Jing, Pengfei, et al.
Veröffentlicht: (2024)
von: Jing, Pengfei, et al.
Veröffentlicht: (2024)
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025)
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025)
Frontier AI's Impact on the Cybersecurity Landscape
von: Potter, Yujin, et al.
Veröffentlicht: (2025)
von: Potter, Yujin, et al.
Veröffentlicht: (2025)
Explainable but Vulnerable: Adversarial Attacks on XAI Explanation in Cybersecurity Applications
von: Mia, Maraz, et al.
Veröffentlicht: (2025)
von: Mia, Maraz, et al.
Veröffentlicht: (2025)
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2024)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2024)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2024)
von: Tihanyi, Norbert, et al.
Veröffentlicht: (2024)
Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
von: Jaffal, Niveen O., et al.
Veröffentlicht: (2025)
von: Jaffal, Niveen O., et al.
Veröffentlicht: (2025)
Are Enterprises Ready for Quantum-Safe Cybersecurity?
von: Le, Tran Duc, et al.
Veröffentlicht: (2025)
von: Le, Tran Duc, et al.
Veröffentlicht: (2025)
Weaponizing Language Models for Cybersecurity Offensive Operations: Automating Vulnerability Assessment Report Validation; A Review Paper
von: Almuhaidib, Abdulrahman S, et al.
Veröffentlicht: (2025)
von: Almuhaidib, Abdulrahman S, et al.
Veröffentlicht: (2025)
A Systematic Approach to Predict the Impact of Cybersecurity Vulnerabilities Using LLMs
von: Høst, Anders Mølmen, et al.
Veröffentlicht: (2025)
von: Høst, Anders Mølmen, et al.
Veröffentlicht: (2025)
CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge
von: Keppler, Gustav, et al.
Veröffentlicht: (2026)
von: Keppler, Gustav, et al.
Veröffentlicht: (2026)
When LLMs Meet Cybersecurity: A Systematic Literature Review
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
von: Lee, Seunghyun, et al.
Veröffentlicht: (2026)
von: Lee, Seunghyun, et al.
Veröffentlicht: (2026)
The New Frontier of Cybersecurity: Emerging Threats and Innovations
von: Dave, Daksh, et al.
Veröffentlicht: (2023)
von: Dave, Daksh, et al.
Veröffentlicht: (2023)
CAI: An Open, Bug Bounty-Ready Cybersecurity AI
von: Mayoral-Vilches, Víctor, et al.
Veröffentlicht: (2025)
von: Mayoral-Vilches, Víctor, et al.
Veröffentlicht: (2025)
SECURE: Benchmarking Large Language Models for Cybersecurity
von: Bhusal, Dipkamal, et al.
Veröffentlicht: (2024)
von: Bhusal, Dipkamal, et al.
Veröffentlicht: (2024)
Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards
von: Hammadia, Taha, et al.
Veröffentlicht: (2026)
von: Hammadia, Taha, et al.
Veröffentlicht: (2026)
Jailbreaking Frontier Foundation Models Through Intention Deception
von: Wang, Xinhe, et al.
Veröffentlicht: (2026)
von: Wang, Xinhe, et al.
Veröffentlicht: (2026)
LLMpatronous: Harnessing the Power of LLMs For Vulnerability Detection
von: Yarra, Rajesh
Veröffentlicht: (2025)
von: Yarra, Rajesh
Veröffentlicht: (2025)
Beyond Single Bugs: Benchmarking Large Language Models for Multi-Vulnerability Detection
von: Pushkar, Chinmay, et al.
Veröffentlicht: (2025)
von: Pushkar, Chinmay, et al.
Veröffentlicht: (2025)
A Survey of Large Language Models in Cybersecurity
von: da Silva, Gabriel de Jesus Coelho, et al.
Veröffentlicht: (2024)
von: da Silva, Gabriel de Jesus Coelho, et al.
Veröffentlicht: (2024)
CurricuLLM: Designing Personalized and Workforce-Aligned Cybersecurity Curricula Using Fine-Tuned LLMs
von: Nijdam, Arthur, et al.
Veröffentlicht: (2026)
von: Nijdam, Arthur, et al.
Veröffentlicht: (2026)
On the Effectiveness of Instruction-Tuning Local LLMs for Identifying Software Vulnerabilities
von: Park, Sangryu, et al.
Veröffentlicht: (2025)
von: Park, Sangryu, et al.
Veröffentlicht: (2025)
Logic Meets Magic: LLMs Cracking Smart Contract Vulnerabilities
von: Xiao, ZeKe, et al.
Veröffentlicht: (2025)
von: Xiao, ZeKe, et al.
Veröffentlicht: (2025)
Enhancing Reverse Engineering: Investigating and Benchmarking Large Language Models for Vulnerability Analysis in Decompiled Binaries
von: Manuel, Dylan, et al.
Veröffentlicht: (2024)
von: Manuel, Dylan, et al.
Veröffentlicht: (2024)
From Texts to Shields: Convergence of Large Language Models and Cybersecurity
von: Li, Tao, et al.
Veröffentlicht: (2025)
von: Li, Tao, et al.
Veröffentlicht: (2025)
Integrative Approaches in Cybersecurity and AI
von: Omar, Marwan
Veröffentlicht: (2024)
von: Omar, Marwan
Veröffentlicht: (2024)
Improving Discovery of Known Software Vulnerability For Enhanced Cybersecurity
von: Sawant, Devesh, et al.
Veröffentlicht: (2024)
von: Sawant, Devesh, et al.
Veröffentlicht: (2024)
Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents
von: Sanz-Gómez, María, et al.
Veröffentlicht: (2025)
von: Sanz-Gómez, María, et al.
Veröffentlicht: (2025)
Swallowing the Poison Pills: Insights from Vulnerability Disparity Among LLMs
von: Yifeng, Peng, et al.
Veröffentlicht: (2025)
von: Yifeng, Peng, et al.
Veröffentlicht: (2025)
LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights
von: Sheng, Ze, et al.
Veröffentlicht: (2025)
von: Sheng, Ze, et al.
Veröffentlicht: (2025)
SecureFalcon: Are We There Yet in Automated Software Vulnerability Detection with LLMs?
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2023)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2023)
Quantifying CBRN Risk in Frontier Models
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2025)
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2025)
Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities
von: Samancioglu, Atil
Veröffentlicht: (2025)
von: Samancioglu, Atil
Veröffentlicht: (2025)
Crimson: Empowering Strategic Reasoning in Cybersecurity through Large Language Models
von: Jin, Jiandong, et al.
Veröffentlicht: (2024)
von: Jin, Jiandong, et al.
Veröffentlicht: (2024)
DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization
von: Tang, Wenxin, et al.
Veröffentlicht: (2026)
von: Tang, Wenxin, et al.
Veröffentlicht: (2026)
Collaborative Intelligence: Topic Modelling of Large Language Model use in Live Cybersecurity Operations
von: Lochner, Martin, et al.
Veröffentlicht: (2025)
von: Lochner, Martin, et al.
Veröffentlicht: (2025)
BugWhisperer: Fine-Tuning LLMs for SoC Hardware Vulnerability Detection
von: Tarek, Shams, et al.
Veröffentlicht: (2025)
von: Tarek, Shams, et al.
Veröffentlicht: (2025)
VADER: A Human-Evaluated Benchmark for Vulnerability Assessment, Detection, Explanation, and Remediation
von: Liu, Ethan TS., et al.
Veröffentlicht: (2025)
von: Liu, Ethan TS., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
von: Jing, Pengfei, et al.
Veröffentlicht: (2024) -
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025) -
Frontier AI's Impact on the Cybersecurity Landscape
von: Potter, Yujin, et al.
Veröffentlicht: (2025) -
Explainable but Vulnerable: Adversarial Attacks on XAI Explanation in Cybersecurity Applications
von: Mia, Maraz, et al.
Veröffentlicht: (2025) -
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2024)