JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongyi, Ye, Jiawei, Wu, Jie, Yan, Tianjie, Wang, Chu, Li, Zhixin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
by: Chen, Xuan, et al.
Published: (2024)
by: Chen, Xuan, et al.
Published: (2024)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
by: Liu, Shuyuan, et al.
Published: (2025)
by: Liu, Shuyuan, et al.
Published: (2025)
ShallowJail: Steering Jailbreaks against Large Language Models
by: Liu, Shang, et al.
Published: (2026)
by: Liu, Shang, et al.
Published: (2026)
Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
by: Xie, Zhixin, et al.
Published: (2025)
by: Xie, Zhixin, et al.
Published: (2025)
JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model
by: Nian, Yi, et al.
Published: (2025)
by: Nian, Yi, et al.
Published: (2025)
TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning
by: Xie, Zhixin, et al.
Published: (2026)
by: Xie, Zhixin, et al.
Published: (2026)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
by: Wang, Xinkai, et al.
Published: (2025)
by: Wang, Xinkai, et al.
Published: (2025)
Activation Surgery: Jailbreaking White-box LLMs without Touching the Prompt
by: Jenny, Maël, et al.
Published: (2026)
by: Jenny, Maël, et al.
Published: (2026)
JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
by: Zhang, Xiaoyu, et al.
Published: (2023)
by: Zhang, Xiaoyu, et al.
Published: (2023)
Black-box Membership Inference Attacks against Fine-tuned Diffusion Models
by: Pang, Yan, et al.
Published: (2023)
by: Pang, Yan, et al.
Published: (2023)
NSmark: Null Space Based Black-box Watermarking Defense Framework for Language Models
by: Zhao, Haodong, et al.
Published: (2024)
by: Zhao, Haodong, et al.
Published: (2024)
TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
by: Wang, Yanting, et al.
Published: (2025)
by: Wang, Yanting, et al.
Published: (2025)
Query Provenance Analysis: Efficient and Robust Defense against Query-based Black-box Attacks
by: Li, Shaofei, et al.
Published: (2024)
by: Li, Shaofei, et al.
Published: (2024)
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
by: Zhang, Xinkai, et al.
Published: (2026)
by: Zhang, Xinkai, et al.
Published: (2026)
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
by: Yi, Biao, et al.
Published: (2025)
by: Yi, Biao, et al.
Published: (2025)
MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
by: Chen, Tailun, et al.
Published: (2025)
by: Chen, Tailun, et al.
Published: (2025)
One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
by: Li, Linbao, et al.
Published: (2025)
by: Li, Linbao, et al.
Published: (2025)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
by: Luo, Weidi, et al.
Published: (2024)
by: Luo, Weidi, et al.
Published: (2024)
JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
by: Chu, Junjie, et al.
Published: (2025)
by: Chu, Junjie, et al.
Published: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
by: Chen, Yunhao, et al.
Published: (2025)
by: Chen, Yunhao, et al.
Published: (2025)
Black-box Optimization of LLM Outputs by Asking for Directions
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
by: Li, Xiao, et al.
Published: (2024)
by: Li, Xiao, et al.
Published: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024)
by: Xiong, Chen, et al.
Published: (2024)
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
by: Wang, Xunguang, et al.
Published: (2024)
by: Wang, Xunguang, et al.
Published: (2024)
Compositional Jailbreaking: An Empirical Analysis of Mutator Chain Interactions in Aligned LLMs
by: Bugnot, Reinelle Jan, et al.
Published: (2026)
by: Bugnot, Reinelle Jan, et al.
Published: (2026)
OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing
by: Chen, Jianming, et al.
Published: (2026)
by: Chen, Jianming, et al.
Published: (2026)
MacPrompt: Maraconic-guided Jailbreak against Text-to-Image Models
by: Ye, Xi, et al.
Published: (2026)
by: Ye, Xi, et al.
Published: (2026)
White-box Membership Inference Attacks against Diffusion Models
by: Pang, Yan, et al.
Published: (2023)
by: Pang, Yan, et al.
Published: (2023)
Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries
by: Wang, Wenqiang, et al.
Published: (2025)
by: Wang, Wenqiang, et al.
Published: (2025)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
by: Wang, Youze, et al.
Published: (2025)
by: Wang, Youze, et al.
Published: (2025)
Mitigating Jailbreaks with Intent-Aware LLMs
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
by: Li, Yanjie, et al.
Published: (2025)
by: Li, Yanjie, et al.
Published: (2025)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
by: Shen, Guangyu, et al.
Published: (2024)
by: Shen, Guangyu, et al.
Published: (2024)
Detecting Fraudulent Services on Quantum Cloud Platforms via Dynamic Fingerprinting
by: Wu, Jindi, et al.
Published: (2024)
by: Wu, Jindi, et al.
Published: (2024)
LLMPirate: LLMs for Black-box Hardware IP Piracy
by: Gohil, Vasudev, et al.
Published: (2024)
by: Gohil, Vasudev, et al.
Published: (2024)
Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
Similar Items
-
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
by: Chen, Xuan, et al.
Published: (2024) -
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
by: Liu, Shuyuan, et al.
Published: (2025) -
ShallowJail: Steering Jailbreaks against Large Language Models
by: Liu, Shang, et al.
Published: (2026) -
Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
by: Xie, Zhixin, et al.
Published: (2025) -
JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model
by: Nian, Yi, et al.
Published: (2025)