Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiaoqun, Liang, Jiacheng, Tang, Luoxi, Ye, Muchao, Ma, Weicheng, Xi, Zhaohan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robustifying Safety-Aligned Large Language Models through Clean Data Curation
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment
by: Tang, Luoxi, et al.
Published: (2025)
by: Tang, Luoxi, et al.
Published: (2025)
SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use
by: Desai, Pratyush, et al.
Published: (2026)
by: Desai, Pratyush, et al.
Published: (2026)
Smart Privacy Policy Assistant: An LLM-Powered System for Transparent and Actionable Privacy Notices
by: Kalvakuntla, Sriharshini, et al.
Published: (2026)
by: Kalvakuntla, Sriharshini, et al.
Published: (2026)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
by: Tong, Haibo, et al.
Published: (2025)
by: Tong, Haibo, et al.
Published: (2025)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
by: Shang, Zhengchun, et al.
Published: (2025)
by: Shang, Zhengchun, et al.
Published: (2025)
Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
by: Meng, Yuqiao, et al.
Published: (2025)
by: Meng, Yuqiao, et al.
Published: (2025)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
by: Jaiswal, Piyush, et al.
Published: (2026)
by: Jaiswal, Piyush, et al.
Published: (2026)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence
by: Meng, Yuqiao, et al.
Published: (2025)
by: Meng, Yuqiao, et al.
Published: (2025)
CyLens: Towards Reinventing Cyber Threat Intelligence in the Paradigm of Agentic Large Language Models
by: Liu, Xiaoqun, et al.
Published: (2025)
by: Liu, Xiaoqun, et al.
Published: (2025)
GUARD-SLM: Token Activation-Based Defense Against Jailbreak Attacks for Small Language Models
by: Mia, Md Jueal, et al.
Published: (2026)
by: Mia, Md Jueal, et al.
Published: (2026)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
by: Li, Xiaohu, et al.
Published: (2025)
by: Li, Xiaohu, et al.
Published: (2025)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
MirrorShield: Towards Universal Defense Against Jailbreaks via Entropy-Guided Mirror Crafting
by: Pu, Rui, et al.
Published: (2025)
by: Pu, Rui, et al.
Published: (2025)
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
by: Yang, Xianglin, et al.
Published: (2025)
by: Yang, Xianglin, et al.
Published: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
by: Xu, Zihao, et al.
Published: (2024)
by: Xu, Zihao, et al.
Published: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
by: Yi, Sibo, et al.
Published: (2024)
by: Yi, Sibo, et al.
Published: (2024)
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
by: Yan, Dong, et al.
Published: (2026)
by: Yan, Dong, et al.
Published: (2026)
DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
by: Cui, Tiehan, et al.
Published: (2025)
by: Cui, Tiehan, et al.
Published: (2025)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
Defenses Against Prompt Attacks Learn Surface Heuristics
by: Li, Shawn, et al.
Published: (2026)
by: Li, Shawn, et al.
Published: (2026)
Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
by: Tsmindashvili, Tatia, et al.
Published: (2025)
by: Tsmindashvili, Tatia, et al.
Published: (2025)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
by: Liu, Shuyuan, et al.
Published: (2025)
by: Liu, Shuyuan, et al.
Published: (2025)
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
by: Ma, Siyuan, et al.
Published: (2024)
by: Ma, Siyuan, et al.
Published: (2024)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
by: Xu, Zhao, et al.
Published: (2024)
by: Xu, Zhao, et al.
Published: (2024)
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
by: Yoon, Sangyeon, et al.
Published: (2026)
by: Yoon, Sangyeon, et al.
Published: (2026)
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
by: Wu, Daoyuan, et al.
Published: (2024)
by: Wu, Daoyuan, et al.
Published: (2024)
JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
by: Pu, Rui, et al.
Published: (2024)
by: Pu, Rui, et al.
Published: (2024)
Attacks and Defenses Against LLM Fingerprinting
by: Kurian, Kevin, et al.
Published: (2025)
by: Kurian, Kevin, et al.
Published: (2025)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
by: Yu, Yongcan, et al.
Published: (2025)
by: Yu, Yongcan, et al.
Published: (2025)
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
by: Zhu, Kaijie, et al.
Published: (2025)
by: Zhu, Kaijie, et al.
Published: (2025)
Similar Items
-
Robustifying Safety-Aligned Large Language Models through Clean Data Curation
by: Liu, Xiaoqun, et al.
Published: (2024) -
POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment
by: Tang, Luoxi, et al.
Published: (2025) -
SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use
by: Desai, Pratyush, et al.
Published: (2026) -
Smart Privacy Policy Assistant: An LLM-Powered System for Transparent and Actionable Privacy Notices
by: Kalvakuntla, Sriharshini, et al.
Published: (2026) -
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
by: Tong, Haibo, et al.
Published: (2025)