From nuclear safety to LLM security: Applying non-probabilistic risk management strategies to build safe and secure LLM-powered systems
Fuente:
arXiv
Saved in:
| Main Authors: | Gutfraind, Alexander, Bier, Vicki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
by: Qi, Jinhu, et al.
Published: (2026)
by: Qi, Jinhu, et al.
Published: (2026)
LLM Scalability Risk for Agentic-AI and Model Supply Chain Security
by: Ahi, Kiarash, et al.
Published: (2026)
by: Ahi, Kiarash, et al.
Published: (2026)
Optimisation of cyber insurance coverage with selection of cost effective security controls
by: Uuganbayar, Ganbayar, et al.
Published: (2025)
by: Uuganbayar, Ganbayar, et al.
Published: (2025)
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
by: Yu, Hao, et al.
Published: (2026)
by: Yu, Hao, et al.
Published: (2026)
Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI
by: Hoang, Tien Dat
Published: (2025)
by: Hoang, Tien Dat
Published: (2025)
Send to which account? Evaluation of an LLM-based Scambaiting System
by: Siadati, Hossein, et al.
Published: (2025)
by: Siadati, Hossein, et al.
Published: (2025)
Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
by: Sanna, Arun Chowdary
Published: (2025)
by: Sanna, Arun Chowdary
Published: (2025)
Sensitivity Uncertainty Alignment in Large Language Models
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
by: Gu, Yongtong, et al.
Published: (2026)
by: Gu, Yongtong, et al.
Published: (2026)
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
by: Morales, Jaime, et al.
Published: (2026)
by: Morales, Jaime, et al.
Published: (2026)
Active Authentication via Korean Keystrokes Under Varying LLM Assistance and Cognitive Contexts
by: Roh, Dong Hyun, et al.
Published: (2025)
by: Roh, Dong Hyun, et al.
Published: (2025)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
Towards Low-Latency and Adaptive Ransomware Detection Using Contrastive Learning
by: Pan, Zhixin, et al.
Published: (2025)
by: Pan, Zhixin, et al.
Published: (2025)
Density-aware Sample-specific Attack
by: Wang, Qiyuan, et al.
Published: (2026)
by: Wang, Qiyuan, et al.
Published: (2026)
Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
by: Collu, Matteo Gioele, et al.
Published: (2023)
by: Collu, Matteo Gioele, et al.
Published: (2023)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
by: Zhang, Tian, et al.
Published: (2026)
by: Zhang, Tian, et al.
Published: (2026)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
by: Chakraborty, Amit, et al.
Published: (2025)
by: Chakraborty, Amit, et al.
Published: (2025)
Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
by: Usman, Rana Muhammad
Published: (2026)
by: Usman, Rana Muhammad
Published: (2026)
A High-Recall Cost-Sensitive Machine Learning Framework for Real-Time Online Banking Transaction Fraud Detection
by: R., Karthikeyan V., et al.
Published: (2026)
by: R., Karthikeyan V., et al.
Published: (2026)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
by: Dawson, Ads, et al.
Published: (2025)
by: Dawson, Ads, et al.
Published: (2025)
The Automation Advantage in AI Red Teaming
by: Mulla, Rob, et al.
Published: (2025)
by: Mulla, Rob, et al.
Published: (2025)
Static Attribution of Android Residential Proxy Malware Using Graph Kernels
by: Clark, Peter, et al.
Published: (2026)
by: Clark, Peter, et al.
Published: (2026)
Scalable APT Malware Classification via Parallel Feature Extraction and GPU-Accelerated Learning
by: Subedar, Noah, et al.
Published: (2025)
by: Subedar, Noah, et al.
Published: (2025)
Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges
by: Muzsai, Lajos, et al.
Published: (2025)
by: Muzsai, Lajos, et al.
Published: (2025)
HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
by: Muzsai, Lajos, et al.
Published: (2024)
by: Muzsai, Lajos, et al.
Published: (2024)
Reverse CAPTCHA: Evaluating LLM Susceptibility to Invisible Unicode Instruction Injection
by: Graves, Marcus
Published: (2026)
by: Graves, Marcus
Published: (2026)
Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems
by: Cai, Yicheng, et al.
Published: (2026)
by: Cai, Yicheng, et al.
Published: (2026)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
by: Chona, Alankrit, et al.
Published: (2026)
by: Chona, Alankrit, et al.
Published: (2026)
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Developing a Strong CPS Defender: An Evolutionary Approach
by: Hu, Qingyuan, et al.
Published: (2025)
by: Hu, Qingyuan, et al.
Published: (2025)
Measuring Harmfulness of Computer-Using Agents
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
Countermind: A Multi-Layered Security Architecture for Large Language Models
by: Schwarz, Dominik
Published: (2025)
by: Schwarz, Dominik
Published: (2025)
Incentivizing Secure Software Development: the Role of Voluntary Audit and Liability Waiver
by: Huang, Ziyuan, et al.
Published: (2024)
by: Huang, Ziyuan, et al.
Published: (2024)
Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study
by: Xu, Luyao, et al.
Published: (2026)
by: Xu, Luyao, et al.
Published: (2026)
Architecture-Agnostic Feature Synergy for Universal Defense Against Heterogeneous Generative Threats
by: Zhang, Bingxue, et al.
Published: (2026)
by: Zhang, Bingxue, et al.
Published: (2026)
LLM Detectors Still Fall Short of Real World: Case of LLM-Generated Short News-Like Posts
by: Gameiro, Henrique Da Silva, et al.
Published: (2024)
by: Gameiro, Henrique Da Silva, et al.
Published: (2024)
A Survey on the Security of Long-Term Memory in LLM Agents: Toward Mnemonic Sovereignty
by: Lin, Zehao, et al.
Published: (2026)
by: Lin, Zehao, et al.
Published: (2026)
SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
by: Adebimpe, Adetayo, et al.
Published: (2025)
by: Adebimpe, Adetayo, et al.
Published: (2025)
Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world
by: South, Tobin, et al.
Published: (2025)
by: South, Tobin, et al.
Published: (2025)
Benchmarking Autonomous Agents against Temporal, Spatial, and Semantic Evasions
by: Ma, Jianan, et al.
Published: (2026)
by: Ma, Jianan, et al.
Published: (2026)
Similar Items
-
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
by: Qi, Jinhu, et al.
Published: (2026) -
LLM Scalability Risk for Agentic-AI and Model Supply Chain Security
by: Ahi, Kiarash, et al.
Published: (2026) -
Optimisation of cyber insurance coverage with selection of cost effective security controls
by: Uuganbayar, Ganbayar, et al.
Published: (2025) -
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
by: Yu, Hao, et al.
Published: (2026) -
Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI
by: Hoang, Tien Dat
Published: (2025)