The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Zhiyuan, Gardiner, Joseph, Belguith, Sana |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Steering in the Shadows: Causal Amplification for Activation Space Attacks in Large Language Models
by: Xu, Zhiyuan, et al.
Published: (2025)
by: Xu, Zhiyuan, et al.
Published: (2025)
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
by: Wu, Xiaodong, et al.
Published: (2025)
by: Wu, Xiaodong, et al.
Published: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
SDD: Self-Degraded Defense against Malicious Fine-tuning
by: Chen, Zixuan, et al.
Published: (2025)
by: Chen, Zixuan, et al.
Published: (2025)
Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
by: Truong, Vu Tuan, et al.
Published: (2026)
by: Truong, Vu Tuan, et al.
Published: (2026)
CoT-Guard: Small Models for Strong Monitoring
by: Diwan, Nirav, et al.
Published: (2026)
by: Diwan, Nirav, et al.
Published: (2026)
DeepSeek Robustness Against Semantic-Character Dual-Space Mutated Prompt Injection
by: Ren, Junyu, et al.
Published: (2026)
by: Ren, Junyu, et al.
Published: (2026)
Optimized detection of cyber-attacks on IoT networks via hybrid deep learning models
by: Bensaoud, Ahmed, et al.
Published: (2025)
by: Bensaoud, Ahmed, et al.
Published: (2025)
Adversarial attacks against Modern Vision-Language Models
by: La Torre, Alejandro Paredes
Published: (2026)
by: La Torre, Alejandro Paredes
Published: (2026)
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
by: Parmar, Manojkumar, et al.
Published: (2025)
by: Parmar, Manojkumar, et al.
Published: (2025)
Detecting Adversarial Fine-tuning with Auditing Agents
by: Egler, Sarah, et al.
Published: (2025)
by: Egler, Sarah, et al.
Published: (2025)
On the use of neurosymbolic AI for defending against cyber attacks
by: Grov, Gudmund, et al.
Published: (2024)
by: Grov, Gudmund, et al.
Published: (2024)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
by: Liu, Guozhi, et al.
Published: (2025)
by: Liu, Guozhi, et al.
Published: (2025)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
Synthetic is all you need: removing the auxiliary data assumption for membership inference attacks against synthetic data
by: Guépin, Florent, et al.
Published: (2023)
by: Guépin, Florent, et al.
Published: (2023)
Detection of ransomware attacks using federated learning based on the CNN model
by: Nguyen, Hong-Nhung, et al.
Published: (2024)
by: Nguyen, Hong-Nhung, et al.
Published: (2024)
Revisiting the attacker's knowledge in inference attacks against Searchable Symmetric Encryption
by: Damie, Marc, et al.
Published: (2025)
by: Damie, Marc, et al.
Published: (2025)
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
by: Huang, Zhengxian, et al.
Published: (2026)
by: Huang, Zhengxian, et al.
Published: (2026)
Comparative Analysis Based on DeepSeek, ChatGPT, and Google Gemini: Features, Techniques, Performance, Future Prospects
by: Rahman, Anichur, et al.
Published: (2025)
by: Rahman, Anichur, et al.
Published: (2025)
DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models
by: Xu, Honghui, et al.
Published: (2025)
by: Xu, Honghui, et al.
Published: (2025)
Class-Aware Adaptive Differential Privacy in Deep Learning for Sensor-Based Fall Detection
by: Sana, Joydeb Kumar
Published: (2026)
by: Sana, Joydeb Kumar
Published: (2026)
Black-box Membership Inference Attacks against Fine-tuned Diffusion Models
by: Pang, Yan, et al.
Published: (2023)
by: Pang, Yan, et al.
Published: (2023)
I can't see it but I can Fine-tune it: On Encrypted Fine-tuning of Transformers using Fully Homomorphic Encryption
by: Panzade, Prajwal, et al.
Published: (2024)
by: Panzade, Prajwal, et al.
Published: (2024)
RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs
by: Xu, Zhiyuan, et al.
Published: (2026)
by: Xu, Zhiyuan, et al.
Published: (2026)
ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
Case Study: Fine-tuning Small Language Models for Accurate and Private CWE Detection in Python Code
by: Bappy, Md. Azizul Hakim, et al.
Published: (2025)
by: Bappy, Md. Azizul Hakim, et al.
Published: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Prompt Injection attack against LLM-integrated Applications
by: Liu, Yi, et al.
Published: (2023)
by: Liu, Yi, et al.
Published: (2023)
Cross-site scripting adversarial attacks based on deep reinforcement learning: Evaluation and extension study
by: Pasini, Samuele, et al.
Published: (2025)
by: Pasini, Samuele, et al.
Published: (2025)
Correlation inference attacks against machine learning models
by: Creţu, Ana-Maria, et al.
Published: (2021)
by: Creţu, Ana-Maria, et al.
Published: (2021)
MEASER: Malware embedding attacks on open-source LLMs
by: Tan, Ming, et al.
Published: (2025)
by: Tan, Ming, et al.
Published: (2025)
DAIRE: A lightweight AI model for real-time detection of Controller Area Network attacks in the Internet of Vehicles
by: Alam, Shahid, et al.
Published: (2026)
by: Alam, Shahid, et al.
Published: (2026)
Token-level Data Selection for Safe LLM Fine-tuning
by: Li, Yanping, et al.
Published: (2026)
by: Li, Yanping, et al.
Published: (2026)
Defending against Stegomalware in Deep Neural Networks with Permutation Symmetry
by: Torpmann-Hagen, Birk, et al.
Published: (2025)
by: Torpmann-Hagen, Birk, et al.
Published: (2025)
Analysis of the vulnerability of machine learning regression models to adversarial attacks using data from 5G wireless networks
by: Legashev, Leonid, et al.
Published: (2025)
by: Legashev, Leonid, et al.
Published: (2025)
CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security LLM Fine-tuning
by: ElZemity, Adel, et al.
Published: (2025)
by: ElZemity, Adel, et al.
Published: (2025)
Visual CoT Makes VLMs Smarter but More Fragile
by: Xu, Chunxue, et al.
Published: (2025)
by: Xu, Chunxue, et al.
Published: (2025)
Context manipulation attacks : Web agents are susceptible to corrupted memory
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
Similar Items
-
Steering in the Shadows: Causal Amplification for Activation Space Attacks in Large Language Models
by: Xu, Zhiyuan, et al.
Published: (2025) -
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
by: Wu, Xiaodong, et al.
Published: (2025) -
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
by: Ying, Zonghao, et al.
Published: (2025) -
SDD: Self-Degraded Defense against Malicious Fine-tuning
by: Chen, Zixuan, et al.
Published: (2025) -
Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
by: Huang, Tiansheng, et al.
Published: (2024)