Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Rongwu, Li, Xiaojian, Chen, Shuo, Xu, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
by: Xu, Rongwu, et al.
Published: (2023)
by: Xu, Rongwu, et al.
Published: (2023)
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
by: Wang, Huandong, et al.
Published: (2025)
by: Wang, Huandong, et al.
Published: (2025)
How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
by: Li, Yuxuan, et al.
Published: (2026)
by: Li, Yuxuan, et al.
Published: (2026)
Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content?
by: Xu, Naen, et al.
Published: (2025)
by: Xu, Naen, et al.
Published: (2025)
On the Suitability of LLM-Driven Agents for Dark Pattern Audits
by: Sun, Chen, et al.
Published: (2026)
by: Sun, Chen, et al.
Published: (2026)
When Agents "Misremember" Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems
by: Xu, Naen, et al.
Published: (2026)
by: Xu, Naen, et al.
Published: (2026)
AI Awareness
by: Li, Xiaojian, et al.
Published: (2025)
by: Li, Xiaojian, et al.
Published: (2025)
Can LLMs Infer Conversational Agent Users' Personality Traits from Chat History?
by: Cögendez, Derya, et al.
Published: (2026)
by: Cögendez, Derya, et al.
Published: (2026)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
by: Yang, Wenkai, et al.
Published: (2024)
by: Yang, Wenkai, et al.
Published: (2024)
Attacks on Third-Party APIs of Large Language Models
by: Zhao, Wanru, et al.
Published: (2024)
by: Zhao, Wanru, et al.
Published: (2024)
Searching for Privacy Risks in LLM Agents via Simulation
by: Zhang, Yanzhe, et al.
Published: (2025)
by: Zhang, Yanzhe, et al.
Published: (2025)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
by: Wang, Kun, et al.
Published: (2025)
by: Wang, Kun, et al.
Published: (2025)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
by: Xu, Huiyu, et al.
Published: (2024)
by: Xu, Huiyu, et al.
Published: (2024)
LATTICE: Evaluating Decision Support Utility of Crypto Agents
by: Chan, Aaron, et al.
Published: (2026)
by: Chan, Aaron, et al.
Published: (2026)
Shadows in the Code: Exploring the Risks and Defenses of LLM-based Multi-Agent Software Development Systems
by: Wang, Xiaoqing, et al.
Published: (2025)
by: Wang, Xiaoqing, et al.
Published: (2025)
Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
by: Luo, Zhifan, et al.
Published: (2025)
by: Luo, Zhifan, et al.
Published: (2025)
SWAN: Semantic Watermarking with Abstract Meaning Representation
by: Ye, Ziping, et al.
Published: (2026)
by: Ye, Ziping, et al.
Published: (2026)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
TombRaider: Entering the Vault of History to Jailbreak Large Language Models
by: Ding, Junchen, et al.
Published: (2025)
by: Ding, Junchen, et al.
Published: (2025)
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
by: Yang, Tiankai, et al.
Published: (2026)
by: Yang, Tiankai, et al.
Published: (2026)
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Large Reasoning Models Are Autonomous Jailbreak Agents
by: Hagendorff, Thilo, et al.
Published: (2025)
by: Hagendorff, Thilo, et al.
Published: (2025)
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
Enforcing Cybersecurity Constraints for LLM-driven Robot Agents for Online Transactions
by: Shah, Shraddha Pradipbhai, et al.
Published: (2025)
by: Shah, Shraddha Pradipbhai, et al.
Published: (2025)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
Phare: A Safety Probe for Large Language Models
by: Jeune, Pierre Le, et al.
Published: (2025)
by: Jeune, Pierre Le, et al.
Published: (2025)
Beyond Context: Large Language Models' Failure to Grasp Users' Intent
by: Hussain, Ahmed M., et al.
Published: (2025)
by: Hussain, Ahmed M., et al.
Published: (2025)
Complete Evasion, Zero Modification: PDF Attacks on AI Text Detection
by: Creo, Aldan
Published: (2025)
by: Creo, Aldan
Published: (2025)
RealHarm: A Collection of Real-World Language Model Application Failures
by: Jeune, Pierre Le, et al.
Published: (2025)
by: Jeune, Pierre Le, et al.
Published: (2025)
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
Blueprints of Trust: AI System Cards for End to End Transparency and Governance
by: Sidhpurwala, Huzaifa, et al.
Published: (2025)
by: Sidhpurwala, Huzaifa, et al.
Published: (2025)
Children's Voice Privacy: First Steps And Emerging Challenges
by: Kulkarni, Ajinkya, et al.
Published: (2025)
by: Kulkarni, Ajinkya, et al.
Published: (2025)
AuditGPT: Auditing Smart Contracts with ChatGPT
by: Xia, Shihao, et al.
Published: (2024)
by: Xia, Shihao, et al.
Published: (2024)
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
by: Lee, Michael S., et al.
Published: (2026)
by: Lee, Michael S., et al.
Published: (2026)
SATA: A Paradigm for LLM Jailbreak via Simple Assistive Task Linkage
by: Dong, Xiaoning, et al.
Published: (2024)
by: Dong, Xiaoning, et al.
Published: (2024)
LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
by: Zhang, Chen Bo Calvin, et al.
Published: (2026)
by: Zhang, Chen Bo Calvin, et al.
Published: (2026)
Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation
by: Shao, Shuo, et al.
Published: (2025)
by: Shao, Shuo, et al.
Published: (2025)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
by: Zhang, Andy K., et al.
Published: (2024)
by: Zhang, Andy K., et al.
Published: (2024)
Similar Items
-
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024) -
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
by: Xu, Rongwu, et al.
Published: (2023) -
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
by: Xu, Rongwu, et al.
Published: (2024) -
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
by: Wang, Huandong, et al.
Published: (2025) -
How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
by: Li, Yuxuan, et al.
Published: (2026)