SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Xiaodong, Li, Xiangman, Li, Qi, Liu, Lingshuang, Ni, Jianbing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Robustness of Watermarking on Text-to-Image Diffusion Models
por: Wu, Xiaodong, et al.
Publicado: (2024)
por: Wu, Xiaodong, et al.
Publicado: (2024)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
por: Wang, Xunguang, et al.
Publicado: (2025)
por: Wang, Xunguang, et al.
Publicado: (2025)
SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
por: Zhao, Wei, et al.
Publicado: (2025)
por: Zhao, Wei, et al.
Publicado: (2025)
Security in LLM-as-a-Judge: A Comprehensive SoK
por: Masoud, Aiman Al, et al.
Publicado: (2026)
por: Masoud, Aiman Al, et al.
Publicado: (2026)
SoK: Robustness in Large Language Models against Jailbreak Attacks
por: Xu, Feiyue, et al.
Publicado: (2026)
por: Xu, Feiyue, et al.
Publicado: (2026)
SoK: Semantic Privacy in Large Language Models
por: Ma, Baihe, et al.
Publicado: (2025)
por: Ma, Baihe, et al.
Publicado: (2025)
SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts
por: Wu, Xiaodong, et al.
Publicado: (2025)
por: Wu, Xiaodong, et al.
Publicado: (2025)
SoK: Security Analysis of Blockchain-based Cryptocurrency
por: Liu, Zekai, et al.
Publicado: (2025)
por: Liu, Zekai, et al.
Publicado: (2025)
PDLRecover: Privacy-preserving Decentralized Model Recovery with Machine Unlearning
por: Li, Xiangman, et al.
Publicado: (2025)
por: Li, Xiangman, et al.
Publicado: (2025)
When There Is No Decoder: Removing Watermarks from Stable Diffusion Models in a No-box Setting
por: Wu, Xiaodong, et al.
Publicado: (2025)
por: Wu, Xiaodong, et al.
Publicado: (2025)
SoK: On the Semantic AI Security in Autonomous Driving
por: Shen, Junjie, et al.
Publicado: (2022)
por: Shen, Junjie, et al.
Publicado: (2022)
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI
por: Zhang, Zelin, et al.
Publicado: (2026)
por: Zhang, Zelin, et al.
Publicado: (2026)
SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
por: Hong, Hanbin, et al.
Publicado: (2025)
por: Hong, Hanbin, et al.
Publicado: (2025)
SoK: Security and Privacy of AI Agents for Blockchain
por: Romandini, Nicolò, et al.
Publicado: (2025)
por: Romandini, Nicolò, et al.
Publicado: (2025)
SoK: Towards Security and Safety of Edge AI
por: Wingarz, Tatjana, et al.
Publicado: (2024)
por: Wingarz, Tatjana, et al.
Publicado: (2024)
SoK: Understanding (New) Security Issues Across AI4Code Use Cases
por: Wu, Qilong, et al.
Publicado: (2025)
por: Wu, Qilong, et al.
Publicado: (2025)
MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech
por: Jin, Yutong, et al.
Publicado: (2026)
por: Jin, Yutong, et al.
Publicado: (2026)
SoK: An Introspective Analysis of RPKI Security
por: Mirdita, Donika, et al.
Publicado: (2024)
por: Mirdita, Donika, et al.
Publicado: (2024)
SoK: Security and Privacy Risks of Healthcare AI
por: Chang, Yuanhaur, et al.
Publicado: (2024)
por: Chang, Yuanhaur, et al.
Publicado: (2024)
SoK: On Gradient Leakage in Federated Learning
por: Du, Jiacheng, et al.
Publicado: (2024)
por: Du, Jiacheng, et al.
Publicado: (2024)
SoK: Security of Programmable Logic Controllers
por: López-Morales, Efrén, et al.
Publicado: (2024)
por: López-Morales, Efrén, et al.
Publicado: (2024)
SoK: Unlearnability and Unlearning for Model Dememorization
por: Zhang, Mengying, et al.
Publicado: (2026)
por: Zhang, Mengying, et al.
Publicado: (2026)
SoK: Trust-Authorization Mismatch in LLM Agent Interactions
por: Shi, Guanquan, et al.
Publicado: (2025)
por: Shi, Guanquan, et al.
Publicado: (2025)
SoK: The Last Line of Defense: On Backdoor Defense Evaluation
por: Abad, Gorka, et al.
Publicado: (2025)
por: Abad, Gorka, et al.
Publicado: (2025)
SoK: Security of EMV Contactless Payment Systems
por: Nezhad, Mahshid Mehr, et al.
Publicado: (2025)
por: Nezhad, Mahshid Mehr, et al.
Publicado: (2025)
SoK: Watermarking for AI-Generated Content
por: Zhao, Xuandong, et al.
Publicado: (2024)
por: Zhao, Xuandong, et al.
Publicado: (2024)
SoK: The Security-Safety Continuum of Multimodal Foundation Models through Information Flow and Global Game-Theoretic Analysis of Asymmetric Threats
por: Sun, Ruoxi, et al.
Publicado: (2024)
por: Sun, Ruoxi, et al.
Publicado: (2024)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
por: Zhang, Heyi, et al.
Publicado: (2025)
por: Zhang, Heyi, et al.
Publicado: (2025)
SoK: Security of Autonomous LLM Agents in Agentic Commerce
por: Mao, Qian'ang, et al.
Publicado: (2026)
por: Mao, Qian'ang, et al.
Publicado: (2026)
SoK: Large Language Model Copyright Auditing via Fingerprinting
por: Shao, Shuo, et al.
Publicado: (2025)
por: Shao, Shuo, et al.
Publicado: (2025)
SoK: Public Blockchain Sharding
por: Barat, Md Mohaimin Al, et al.
Publicado: (2024)
por: Barat, Md Mohaimin Al, et al.
Publicado: (2024)
SoK: Analysis techniques for WebAssembly
por: Harnes, Håkon, et al.
Publicado: (2024)
por: Harnes, Håkon, et al.
Publicado: (2024)
SoK: Leveraging Transformers for Malware Analysis
por: Kunwar, Pradip, et al.
Publicado: (2024)
por: Kunwar, Pradip, et al.
Publicado: (2024)
SoK: Evolution, Security, and Fundamental Properties of Transactional Systems
por: Waterpeace, Sky Pelletier, et al.
Publicado: (2026)
por: Waterpeace, Sky Pelletier, et al.
Publicado: (2026)
SoK: Verifiable Cross-Silo FL
por: Korneev, Aleksei, et al.
Publicado: (2024)
por: Korneev, Aleksei, et al.
Publicado: (2024)
SoK: Prompt Hacking of Large Language Models
por: Rababah, Baha, et al.
Publicado: (2024)
por: Rababah, Baha, et al.
Publicado: (2024)
SoK: Runtime Integrity
por: Ammar, Mahmoud, et al.
Publicado: (2024)
por: Ammar, Mahmoud, et al.
Publicado: (2024)
SoK: How Robust is Audio Watermarking in Generative AI models?
por: Wen, Yizhu, et al.
Publicado: (2025)
por: Wen, Yizhu, et al.
Publicado: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
SoK: Speedy Secure Finality
por: Saraswat, Yash, et al.
Publicado: (2025)
por: Saraswat, Yash, et al.
Publicado: (2025)
Ejemplares similares
-
Robustness of Watermarking on Text-to-Image Diffusion Models
por: Wu, Xiaodong, et al.
Publicado: (2024) -
SoK: Evaluating Jailbreak Guardrails for Large Language Models
por: Wang, Xunguang, et al.
Publicado: (2025) -
SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
por: Zhao, Wei, et al.
Publicado: (2025) -
Security in LLM-as-a-Judge: A Comprehensive SoK
por: Masoud, Aiman Al, et al.
Publicado: (2026) -
SoK: Robustness in Large Language Models against Jailbreak Attacks
por: Xu, Feiyue, et al.
Publicado: (2026)