BSPA: Exploring Black-box Stealthy Prompt Attacks against Image Generators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tian, Yu, Yang, Xiao, Dong, Yinpeng, Yang, Heming, Su, Hang, Zhu, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
von: Yang, Yong, et al.
Veröffentlicht: (2024)
von: Yang, Yong, et al.
Veröffentlicht: (2024)
Efficient Black-box Adversarial Attacks via Bayesian Optimization Guided by a Function Prior
von: Cheng, Shuyu, et al.
Veröffentlicht: (2024)
von: Cheng, Shuyu, et al.
Veröffentlicht: (2024)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
von: Huang, Yao, et al.
Veröffentlicht: (2025)
von: Huang, Yao, et al.
Veröffentlicht: (2025)
Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings
von: Zhang, Yuanhe, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhe, et al.
Veröffentlicht: (2024)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
von: Li, Hao, et al.
Veröffentlicht: (2026)
von: Li, Hao, et al.
Veröffentlicht: (2026)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
DiscourseFlip: An Oblique Discourse-Level Opinion Manipulation Attack against Black-box Retrieval-Augmented Generation
von: Gong, Yuyang, et al.
Veröffentlicht: (2026)
von: Gong, Yuyang, et al.
Veröffentlicht: (2026)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
Prompt Stealing Attacks Against Large Language Models
von: Sha, Zeyang, et al.
Veröffentlicht: (2024)
von: Sha, Zeyang, et al.
Veröffentlicht: (2024)
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
von: Yi, Biao, et al.
Veröffentlicht: (2025)
von: Yi, Biao, et al.
Veröffentlicht: (2025)
Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
von: Na, CheolWon, et al.
Veröffentlicht: (2025)
von: Na, CheolWon, et al.
Veröffentlicht: (2025)
One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
von: Li, Linbao, et al.
Veröffentlicht: (2025)
von: Li, Linbao, et al.
Veröffentlicht: (2025)
MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots
von: Wang, Yuhao, et al.
Veröffentlicht: (2026)
von: Wang, Yuhao, et al.
Veröffentlicht: (2026)
Denial-of-Service Poisoning Attacks against Large Language Models
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
Disabling Self-Correction in Retrieval-Augmented Generation via Stealthy Retriever Poisoning
von: Dai, Yanbo, et al.
Veröffentlicht: (2025)
von: Dai, Yanbo, et al.
Veröffentlicht: (2025)
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2024)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
von: Xue, Eric, et al.
Veröffentlicht: (2025)
von: Xue, Eric, et al.
Veröffentlicht: (2025)
Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors
von: Yin, Rui, et al.
Veröffentlicht: (2026)
von: Yin, Rui, et al.
Veröffentlicht: (2026)
T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
von: Zhao, Gejian, et al.
Veröffentlicht: (2025)
von: Zhao, Gejian, et al.
Veröffentlicht: (2025)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
von: Wang, Jiawen, et al.
Veröffentlicht: (2025)
von: Wang, Jiawen, et al.
Veröffentlicht: (2025)
Stealthy Targeted Backdoor Attacks against Image Captioning
von: Fan, Wenshu, et al.
Veröffentlicht: (2024)
von: Fan, Wenshu, et al.
Veröffentlicht: (2024)
Exploiting Class Probabilities for Black-box Sentence-level Attacks
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
Black-Box Guardrail Reverse-engineering Attack
von: Yao, Hongwei, et al.
Veröffentlicht: (2025)
von: Yao, Hongwei, et al.
Veröffentlicht: (2025)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
von: He, Jiaming, et al.
Veröffentlicht: (2024)
von: He, Jiaming, et al.
Veröffentlicht: (2024)
Proactive defense against LLM Jailbreak
von: Zhao, Weiliang, et al.
Veröffentlicht: (2025)
von: Zhao, Weiliang, et al.
Veröffentlicht: (2025)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
von: Li, Xiang, et al.
Veröffentlicht: (2025) -
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
von: Yang, Yong, et al.
Veröffentlicht: (2024) -
Efficient Black-box Adversarial Attacks via Bayesian Optimization Guided by a Function Prior
von: Cheng, Shuyu, et al.
Veröffentlicht: (2024) -
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
von: Wei, Jiali, et al.
Veröffentlicht: (2026) -
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)