LLM Cyber Evaluations Don't Capture Real-World Risk
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lukošiūtė, Kamilė, Swanda, Adam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
von: Hernandez, Adriano
Veröffentlicht: (2024)
von: Hernandez, Adriano
Veröffentlicht: (2024)
RLShield: Practical Multi-Agent RL for Financial Cyber Defense with Attack-Surface MDPs and Real-Time Response Orchestration
von: Nayak, Srikumar
Veröffentlicht: (2026)
von: Nayak, Srikumar
Veröffentlicht: (2026)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
von: Moia, Vitor Hugo Galhardo, et al.
Veröffentlicht: (2025)
von: Moia, Vitor Hugo Galhardo, et al.
Veröffentlicht: (2025)
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
von: Zhang, Andy K., et al.
Veröffentlicht: (2025)
von: Zhang, Andy K., et al.
Veröffentlicht: (2025)
AutoMalDesc: Large-Scale Script Analysis for Cyber Threat Research
von: Apostu, Alexandru-Mihai, et al.
Veröffentlicht: (2025)
von: Apostu, Alexandru-Mihai, et al.
Veröffentlicht: (2025)
AnnoCTR: A Dataset for Detecting and Linking Entities, Tactics, and Techniques in Cyber Threat Reports
von: Lange, Lukas, et al.
Veröffentlicht: (2024)
von: Lange, Lukas, et al.
Veröffentlicht: (2024)
There Are No Silly Questions: Evaluation of Offline LLM Capabilities from a Turkish Perspective
von: Yilmaz, Edibe, et al.
Veröffentlicht: (2026)
von: Yilmaz, Edibe, et al.
Veröffentlicht: (2026)
Clio: Privacy-Preserving Insights into Real-World AI Use
von: Tamkin, Alex, et al.
Veröffentlicht: (2024)
von: Tamkin, Alex, et al.
Veröffentlicht: (2024)
Federated In-Context LLM Agent Learning
von: Wu, Panlong, et al.
Veröffentlicht: (2024)
von: Wu, Panlong, et al.
Veröffentlicht: (2024)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
von: Wu, Yiran, et al.
Veröffentlicht: (2025)
von: Wu, Yiran, et al.
Veröffentlicht: (2025)
Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness
von: Shafee, Samaneh, et al.
Veröffentlicht: (2024)
von: Shafee, Samaneh, et al.
Veröffentlicht: (2024)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
Policy-Invisible Violations in LLM-Based Agents
von: Wu, Jie, et al.
Veröffentlicht: (2026)
von: Wu, Jie, et al.
Veröffentlicht: (2026)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
von: Zhang, Andy K., et al.
Veröffentlicht: (2024)
von: Zhang, Andy K., et al.
Veröffentlicht: (2024)
Efficient LLM Moderation with Multi-Layer Latent Prototypes
von: Chrabąszcz, Maciej, et al.
Veröffentlicht: (2025)
von: Chrabąszcz, Maciej, et al.
Veröffentlicht: (2025)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Adaptive Instruction Composition for Automated LLM Red-Teaming
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
GaussMark: A Practical Approach for Structural Watermarking of Language Models
von: Block, Adam, et al.
Veröffentlicht: (2025)
von: Block, Adam, et al.
Veröffentlicht: (2025)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
The Kerimov-Alekberli Model: An Information-Geometric Framework for Real-Time System Stability
von: Karimov, Hikmat, et al.
Veröffentlicht: (2026)
von: Karimov, Hikmat, et al.
Veröffentlicht: (2026)
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
von: Fonseca, Joao, et al.
Veröffentlicht: (2025)
von: Fonseca, Joao, et al.
Veröffentlicht: (2025)
The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting
von: Zhang, Hanxiu, et al.
Veröffentlicht: (2025)
von: Zhang, Hanxiu, et al.
Veröffentlicht: (2025)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
von: Dobre, David, et al.
Veröffentlicht: (2025)
von: Dobre, David, et al.
Veröffentlicht: (2025)
From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
von: Zhang, Mohan, et al.
Veröffentlicht: (2026) -
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
von: Hernandez, Adriano
Veröffentlicht: (2024) -
RLShield: Practical Multi-Agent RL for Financial Cyber Defense with Attack-Surface MDPs and Real-Time Response Orchestration
von: Nayak, Srikumar
Veröffentlicht: (2026) -
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
von: Wang, Zhun, et al.
Veröffentlicht: (2025) -
LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
von: Moia, Vitor Hugo Galhardo, et al.
Veröffentlicht: (2025)