PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Ruoxi, Ding, Yizhong, Cao, Shuirong, Duan, Ranjie, Jia, Xiaoshuang, Yuan, Shaowei, Qin, Simeng, Wang, Zhiqiang, Jia, Xiaojun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
von: Li, Jianan, et al.
Veröffentlicht: (2026)
von: Li, Jianan, et al.
Veröffentlicht: (2026)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
von: Teng, Ma, et al.
Veröffentlicht: (2024)
von: Teng, Ma, et al.
Veröffentlicht: (2024)
Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
von: Huang, Xun, et al.
Veröffentlicht: (2026)
von: Huang, Xun, et al.
Veröffentlicht: (2026)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2025)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2025)
Gibberish is All You Need for Membership Inference Detection in Contrastive Language-Audio Pretraining
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models
von: Gao, Sensen, et al.
Veröffentlicht: (2024)
von: Gao, Sensen, et al.
Veröffentlicht: (2024)
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
von: Jia, Xiaojun, et al.
Veröffentlicht: (2025)
von: Jia, Xiaojun, et al.
Veröffentlicht: (2025)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
von: Cao, Shuirong, et al.
Veröffentlicht: (2024)
von: Cao, Shuirong, et al.
Veröffentlicht: (2024)
Untargeted Jailbreak Attack
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
von: Zhong, Xingwei, et al.
Veröffentlicht: (2025)
von: Zhong, Xingwei, et al.
Veröffentlicht: (2025)
SGHA-Attack: Semantic-Guided Hierarchical Alignment for Transferable Targeted Attacks on Vision-Language Models
von: Wang, Haobo, et al.
Veröffentlicht: (2026)
von: Wang, Haobo, et al.
Veröffentlicht: (2026)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
von: Gohil, Vasudev
Veröffentlicht: (2025)
von: Gohil, Vasudev
Veröffentlicht: (2025)
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
von: Chen, Kejia, et al.
Veröffentlicht: (2026)
von: Chen, Kejia, et al.
Veröffentlicht: (2026)
Enhancing Jailbreak Attacks with Diversity Guidance
von: Zhang, Xu, et al.
Veröffentlicht: (2024)
von: Zhang, Xu, et al.
Veröffentlicht: (2024)
All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks
von: Takemoto, Kazuhiro
Veröffentlicht: (2024)
von: Takemoto, Kazuhiro
Veröffentlicht: (2024)
$B^4$: A Black-Box Scrubbing Attack on LLM Watermarks
von: Huang, Baizhou, et al.
Veröffentlicht: (2024)
von: Huang, Baizhou, et al.
Veröffentlicht: (2024)
PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving Systems
von: Guo, Qi, et al.
Veröffentlicht: (2025)
von: Guo, Qi, et al.
Veröffentlicht: (2025)
Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language Attack
von: Jia, Xiaojun, et al.
Veröffentlicht: (2024)
von: Jia, Xiaojun, et al.
Veröffentlicht: (2024)
AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery
von: Wang, Haowei, et al.
Veröffentlicht: (2025)
von: Wang, Haowei, et al.
Veröffentlicht: (2025)
Antelope: Potent and Concealed Jailbreak Attack Strategy
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework
von: Ma, Binhao, et al.
Veröffentlicht: (2025)
von: Ma, Binhao, et al.
Veröffentlicht: (2025)
AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
von: Chen, Guangke, et al.
Veröffentlicht: (2025)
von: Chen, Guangke, et al.
Veröffentlicht: (2025)
Strata-Sword: A Hierarchical Safety Evaluation towards LLMs based on Reasoning Complexity of Jailbreak Instructions
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models
von: Yang, Yiqi, et al.
Veröffentlicht: (2024)
von: Yang, Yiqi, et al.
Veröffentlicht: (2024)
SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG
von: Si, Xiaonan, et al.
Veröffentlicht: (2025)
von: Si, Xiaonan, et al.
Veröffentlicht: (2025)
GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations
von: Liu, Xinwei, et al.
Veröffentlicht: (2025)
von: Liu, Xinwei, et al.
Veröffentlicht: (2025)
Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
von: Dabas, Mahavir, et al.
Veröffentlicht: (2025)
von: Dabas, Mahavir, et al.
Veröffentlicht: (2025)
ODAR: Principled Adaptive Routing for LLM Reasoning via Active Inference
von: Ma, Siyuan, et al.
Veröffentlicht: (2026)
von: Ma, Siyuan, et al.
Veröffentlicht: (2026)
Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment
von: Jia, Xiaojun, et al.
Veröffentlicht: (2025)
von: Jia, Xiaojun, et al.
Veröffentlicht: (2025)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
Activation-Guided Local Editing for Jailbreaking Attacks
von: Wang, Jiecong, et al.
Veröffentlicht: (2025)
von: Wang, Jiecong, et al.
Veröffentlicht: (2025)
Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
von: Nurlanov, Zhakshylyk, et al.
Veröffentlicht: (2026)
von: Nurlanov, Zhakshylyk, et al.
Veröffentlicht: (2026)
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2026)
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2026)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
von: Wang, Libo
Veröffentlicht: (2024)
von: Wang, Libo
Veröffentlicht: (2024)
Breaking the Black-Box: Confidence-Guided Model Inversion Attack for Distribution Shift
von: Liu, Xinhao, et al.
Veröffentlicht: (2024)
von: Liu, Xinhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
von: Li, Jianan, et al.
Veröffentlicht: (2026) -
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
von: Teng, Ma, et al.
Veröffentlicht: (2024) -
Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
von: Huang, Xun, et al.
Veröffentlicht: (2026) -
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2025) -
Gibberish is All You Need for Membership Inference Detection in Contrastive Language-Audio Pretraining
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)