Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Parmar, Manojkumar, Govindarajulu, Yuvaraj |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MISLEAD: Manipulating Importance of Selected features for Learning Epsilon in Evasion Attack Deception
von: Khazanchi, Vidit, et al.
Veröffentlicht: (2024)
von: Khazanchi, Vidit, et al.
Veröffentlicht: (2024)
Enhancing TinyML Security: Study of Adversarial Attack Transferability
von: Shah, Parin, et al.
Veröffentlicht: (2024)
von: Shah, Parin, et al.
Veröffentlicht: (2024)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
Decoding Deception: Understanding Automatic Speech Recognition Vulnerabilities in Evasion and Poisoning Attacks
von: G, Aravindhan, et al.
Veröffentlicht: (2025)
von: G, Aravindhan, et al.
Veröffentlicht: (2025)
Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2026)
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2026)
Mapping LLM Security Landscapes: A Comprehensive Stakeholder Risk Assessment Proposal
von: Pankajakshan, Rahul, et al.
Veröffentlicht: (2024)
von: Pankajakshan, Rahul, et al.
Veröffentlicht: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
Intent Laundering: AI Safety Datasets Are Not What They Seem
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
Trustworthy AI: Safety, Bias, and Privacy -- A Survey
von: Fang, Xingli, et al.
Veröffentlicht: (2025)
von: Fang, Xingli, et al.
Veröffentlicht: (2025)
VidModEx: Interpretable and Efficient Black Box Model Extraction for High-Dimensional Spaces
von: Kumar, Somnath Sendhil, et al.
Veröffentlicht: (2024)
von: Kumar, Somnath Sendhil, et al.
Veröffentlicht: (2024)
Lifelong Safety Alignment for Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Differentially Private Deep Model-Based Reinforcement Learning
von: Rio, Alexandre, et al.
Veröffentlicht: (2024)
von: Rio, Alexandre, et al.
Veröffentlicht: (2024)
On the Role of Attention Heads in Large Language Model Safety
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
von: Li, Lijun, et al.
Veröffentlicht: (2024)
von: Li, Lijun, et al.
Veröffentlicht: (2024)
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
von: Fonseca, Joao, et al.
Veröffentlicht: (2025)
von: Fonseca, Joao, et al.
Veröffentlicht: (2025)
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
von: Wu, Xiaodong, et al.
Veröffentlicht: (2025)
von: Wu, Xiaodong, et al.
Veröffentlicht: (2025)
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
von: Ren, Qibing, et al.
Veröffentlicht: (2024)
von: Ren, Qibing, et al.
Veröffentlicht: (2024)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
CPE-Identifier: Automated CPE identification and CVE summaries annotation with Deep Learning and NLP
von: Hu, Wanyu, et al.
Veröffentlicht: (2024)
von: Hu, Wanyu, et al.
Veröffentlicht: (2024)
Evading Deep Learning-Based Malware Detectors via Obfuscation: A Deep Reinforcement Learning Approach
von: Etter, Brian, et al.
Veröffentlicht: (2024)
von: Etter, Brian, et al.
Veröffentlicht: (2024)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
von: Vega, Jason, et al.
Veröffentlicht: (2023)
von: Vega, Jason, et al.
Veröffentlicht: (2023)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
Generative AI Security: Challenges and Countermeasures
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
Mitigating Deep Reinforcement Learning Backdoors in the Neural Activation Space
von: Vyas, Sanyam, et al.
Veröffentlicht: (2024)
von: Vyas, Sanyam, et al.
Veröffentlicht: (2024)
MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs
von: Kan, Chun Yan Ryan, et al.
Veröffentlicht: (2026)
von: Kan, Chun Yan Ryan, et al.
Veröffentlicht: (2026)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
von: He, Luxi, et al.
Veröffentlicht: (2024)
von: He, Luxi, et al.
Veröffentlicht: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
von: Shetty, Anudeex, et al.
Veröffentlicht: (2026)
von: Shetty, Anudeex, et al.
Veröffentlicht: (2026)
DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
AuthorMist: Evading AI Text Detectors with Reinforcement Learning
von: David, Isaac, et al.
Veröffentlicht: (2025)
von: David, Isaac, et al.
Veröffentlicht: (2025)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
von: Wang, Kun, et al.
Veröffentlicht: (2025)
von: Wang, Kun, et al.
Veröffentlicht: (2025)
Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
Deep Reinforcement Learning for Phishing Detection with Transformer-Based Semantic Features
von: Faisal, Aseer Al
Veröffentlicht: (2025)
von: Faisal, Aseer Al
Veröffentlicht: (2025)
Advanced Persistent Threats (APT) Attribution Using Deep Reinforcement Learning
von: Basnet, Animesh Singh, et al.
Veröffentlicht: (2024)
von: Basnet, Animesh Singh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MISLEAD: Manipulating Importance of Selected features for Learning Epsilon in Evasion Attack Deception
von: Khazanchi, Vidit, et al.
Veröffentlicht: (2024) -
Enhancing TinyML Security: Study of Adversarial Attack Transferability
von: Shah, Parin, et al.
Veröffentlicht: (2024) -
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
von: Ying, Zonghao, et al.
Veröffentlicht: (2025) -
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
von: Naseh, Ali, et al.
Veröffentlicht: (2025) -
Decoding Deception: Understanding Automatic Speech Recognition Vulnerabilities in Evasion and Poisoning Attacks
von: G, Aravindhan, et al.
Veröffentlicht: (2025)