Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
Fuente:
arXiv
Salvato in:
| Autori principali: | Ying, Zonghao, Zheng, Guangyi, Huang, Yongxin, Zhang, Deyue, Zhang, Wenxin, Zou, Quanchen, Liu, Aishan, Liu, Xianglong, Tao, Dacheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
di: Liu, Zhe, et al.
Pubblicazione: (2026)
di: Liu, Zhe, et al.
Pubblicazione: (2026)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
Evolving Deception: When Agents Evolve, Deception Wins
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
di: Xu, Wenzhuo, et al.
Pubblicazione: (2026)
di: Xu, Wenzhuo, et al.
Pubblicazione: (2026)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
di: Zou, Quanchen, et al.
Pubblicazione: (2026)
di: Zou, Quanchen, et al.
Pubblicazione: (2026)
Robust Privacy: Inference-Time Privacy through Certified Robustness
di: Jin, Jiankai, et al.
Pubblicazione: (2026)
di: Jin, Jiankai, et al.
Pubblicazione: (2026)
Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application
di: Xu, Wenzhuo, et al.
Pubblicazione: (2025)
di: Xu, Wenzhuo, et al.
Pubblicazione: (2025)
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
di: Zou, Quanchen, et al.
Pubblicazione: (2025)
di: Zou, Quanchen, et al.
Pubblicazione: (2025)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
di: Mu, Junjie, et al.
Pubblicazione: (2025)
di: Mu, Junjie, et al.
Pubblicazione: (2025)
Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
SoK: Understanding Vulnerabilities in the Large Language Model Supply Chain
di: Wang, Shenao, et al.
Pubblicazione: (2025)
di: Wang, Shenao, et al.
Pubblicazione: (2025)
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
di: Yang, Feiyu, et al.
Pubblicazione: (2025)
di: Yang, Feiyu, et al.
Pubblicazione: (2025)
Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
di: Zhang, Deyue, et al.
Pubblicazione: (2025)
di: Zhang, Deyue, et al.
Pubblicazione: (2025)
Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking
di: Chen, Moyang, et al.
Pubblicazione: (2026)
di: Chen, Moyang, et al.
Pubblicazione: (2026)
DeepSeek Robustness Against Semantic-Character Dual-Space Mutated Prompt Injection
di: Ren, Junyu, et al.
Pubblicazione: (2026)
di: Ren, Junyu, et al.
Pubblicazione: (2026)
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
di: Wu, Xiaodong, et al.
Pubblicazione: (2025)
di: Wu, Xiaodong, et al.
Pubblicazione: (2025)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
di: Ren, Zhiyao, et al.
Pubblicazione: (2025)
di: Ren, Zhiyao, et al.
Pubblicazione: (2025)
Understanding Help Seeking for Digital Privacy, Safety, and Security
di: Thomas, Kurt, et al.
Pubblicazione: (2026)
di: Thomas, Kurt, et al.
Pubblicazione: (2026)
DLP: towards active defense against backdoor attacks with decoupled learning process
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
NBA: defensive distillation for backdoor removal via neural behavior alignment
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
di: Parmar, Manojkumar, et al.
Pubblicazione: (2025)
di: Parmar, Manojkumar, et al.
Pubblicazione: (2025)
Comparative Analysis Based on DeepSeek, ChatGPT, and Google Gemini: Features, Techniques, Performance, Future Prospects
di: Rahman, Anichur, et al.
Pubblicazione: (2025)
di: Rahman, Anichur, et al.
Pubblicazione: (2025)
The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models
di: Xu, Zhiyuan, et al.
Pubblicazione: (2025)
di: Xu, Zhiyuan, et al.
Pubblicazione: (2025)
Compromising Embodied Agents with Contextual Backdoor Attacks
di: Liu, Aishan, et al.
Pubblicazione: (2024)
di: Liu, Aishan, et al.
Pubblicazione: (2024)
Demystifying and Detecting Agentic Workflow Injection Vulnerabilities in GitHub Actions
di: Wang, Shenao, et al.
Pubblicazione: (2026)
di: Wang, Shenao, et al.
Pubblicazione: (2026)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
di: Liu, Xuxu, et al.
Pubblicazione: (2025)
di: Liu, Xuxu, et al.
Pubblicazione: (2025)
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models
di: Liang, Siyuan, et al.
Pubblicazione: (2025)
di: Liang, Siyuan, et al.
Pubblicazione: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
di: Naseh, Ali, et al.
Pubblicazione: (2025)
di: Naseh, Ali, et al.
Pubblicazione: (2025)
Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs
di: Zhou, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhou, Yuxuan, et al.
Pubblicazione: (2025)
Exploring Traffic Simulation and Cybersecurity Strategies Using Large Language Models
di: Gao, Lu, et al.
Pubblicazione: (2025)
di: Gao, Lu, et al.
Pubblicazione: (2025)
Exploiting Defenses against GAN-Based Feature Inference Attacks in Federated Learning
di: Luo, Xinjian, et al.
Pubblicazione: (2020)
di: Luo, Xinjian, et al.
Pubblicazione: (2020)
FedMABA: Towards Fair Federated Learning through Multi-Armed Bandits Allocation
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
di: Ye, Mang, et al.
Pubblicazione: (2025)
di: Ye, Mang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
di: Ying, Zonghao, et al.
Pubblicazione: (2024) -
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
di: Ying, Zonghao, et al.
Pubblicazione: (2025) -
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
di: Ying, Zonghao, et al.
Pubblicazione: (2024) -
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
di: Liu, Zhe, et al.
Pubblicazione: (2026) -
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
di: Ying, Zonghao, et al.
Pubblicazione: (2026)