Discovering Universal Semantic Triggers for Text-to-Image Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhai, Shengfang, Wang, Weilong, Li, Jiajun, Dong, Yinpeng, Su, Hang, Shen, Qingni |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood Discrepancy
di: Zhai, Shengfang, et al.
Pubblicazione: (2024)
di: Zhai, Shengfang, et al.
Pubblicazione: (2024)
Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation Variation
di: Zhai, Shengfang, et al.
Pubblicazione: (2025)
di: Zhai, Shengfang, et al.
Pubblicazione: (2025)
Purify Once, Edit Freely: Breaking Image Protections under Model Mismatch
di: Zhao, Qichen, et al.
Pubblicazione: (2026)
di: Zhao, Qichen, et al.
Pubblicazione: (2026)
Life-Cycle Routing Vulnerabilities of LLM Router
di: Lin, Qiqi, et al.
Pubblicazione: (2025)
di: Lin, Qiqi, et al.
Pubblicazione: (2025)
Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots
di: Wang, Yuhao, et al.
Pubblicazione: (2026)
di: Wang, Yuhao, et al.
Pubblicazione: (2026)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
di: Wang, Youze, et al.
Pubblicazione: (2025)
di: Wang, Youze, et al.
Pubblicazione: (2025)
Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models
di: Tian, Zhihua, et al.
Pubblicazione: (2025)
di: Tian, Zhihua, et al.
Pubblicazione: (2025)
DMark: Order-Agnostic Watermarking for Diffusion Large Language Models
di: Wu, Linyu, et al.
Pubblicazione: (2025)
di: Wu, Linyu, et al.
Pubblicazione: (2025)
Semantic Encryption: Secure and Effective Interaction with Cloud-based Large Language Models via Semantic Transformation
di: Chen, Dong, et al.
Pubblicazione: (2025)
di: Chen, Dong, et al.
Pubblicazione: (2025)
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
di: Chen, Tianxin, et al.
Pubblicazione: (2026)
di: Chen, Tianxin, et al.
Pubblicazione: (2026)
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
di: Lin, Chenhao, et al.
Pubblicazione: (2025)
di: Lin, Chenhao, et al.
Pubblicazione: (2025)
Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
di: Liu, Tong, et al.
Pubblicazione: (2024)
di: Liu, Tong, et al.
Pubblicazione: (2024)
Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling
di: Cao, Yichuan, et al.
Pubblicazione: (2025)
di: Cao, Yichuan, et al.
Pubblicazione: (2025)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
di: Liu, Yue, et al.
Pubblicazione: (2025)
di: Liu, Yue, et al.
Pubblicazione: (2025)
IMMACULATE: A Practical LLM Auditing Framework via Verifiable Computation
di: Guo, Yanpei, et al.
Pubblicazione: (2026)
di: Guo, Yanpei, et al.
Pubblicazione: (2026)
When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks
di: Chen, Shenyang, et al.
Pubblicazione: (2026)
di: Chen, Shenyang, et al.
Pubblicazione: (2026)
The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
di: Bullwinkel, Blake, et al.
Pubblicazione: (2026)
di: Bullwinkel, Blake, et al.
Pubblicazione: (2026)
DPImageBench: A Unified Benchmark for Differentially Private Image Synthesis
di: Gong, Chen, et al.
Pubblicazione: (2025)
di: Gong, Chen, et al.
Pubblicazione: (2025)
From Easy to Hard: Building a Shortcut for Differentially Private Image Synthesis
di: Li, Kecen, et al.
Pubblicazione: (2025)
di: Li, Kecen, et al.
Pubblicazione: (2025)
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
di: Liang, Jiashuo, et al.
Pubblicazione: (2024)
di: Liang, Jiashuo, et al.
Pubblicazione: (2024)
Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation
di: Gao, Yilan, et al.
Pubblicazione: (2026)
di: Gao, Yilan, et al.
Pubblicazione: (2026)
Joint Universal Adversarial Perturbations with Interpretations
di: Ning, Liang-bo, et al.
Pubblicazione: (2024)
di: Ning, Liang-bo, et al.
Pubblicazione: (2024)
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
di: Zheng, Jingyi, et al.
Pubblicazione: (2025)
di: Zheng, Jingyi, et al.
Pubblicazione: (2025)
On Google's SynthID-Text LLM Watermarking System: Theoretical Analysis and Empirical Validation
di: Omidi, Romina, et al.
Pubblicazione: (2026)
di: Omidi, Romina, et al.
Pubblicazione: (2026)
Real-world Adversarial Defense against Patch Attacks based on Diffusion Model
di: Wei, Xingxing, et al.
Pubblicazione: (2024)
di: Wei, Xingxing, et al.
Pubblicazione: (2024)
DIFFender: Diffusion-Based Adversarial Defense against Patch Attacks
di: Kang, Caixin, et al.
Pubblicazione: (2023)
di: Kang, Caixin, et al.
Pubblicazione: (2023)
ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users
di: Li, Guanlin, et al.
Pubblicazione: (2024)
di: Li, Guanlin, et al.
Pubblicazione: (2024)
When Convenience Becomes Risk: A Semantic View of Under-Specification in Host-Acting Agents
di: Lu, Di, et al.
Pubblicazione: (2026)
di: Lu, Di, et al.
Pubblicazione: (2026)
Automatic Jailbreaking of the Text-to-Image Generative AI Systems
di: Kim, Minseon, et al.
Pubblicazione: (2024)
di: Kim, Minseon, et al.
Pubblicazione: (2024)
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
di: Huang, Yao, et al.
Pubblicazione: (2025)
di: Huang, Yao, et al.
Pubblicazione: (2025)
Unlink to Unlearn: Simplifying Edge Unlearning in GNNs
di: Tan, Jiajun, et al.
Pubblicazione: (2024)
di: Tan, Jiajun, et al.
Pubblicazione: (2024)
Attack-Resistant Watermarking for AIGC Image Forensics via Diffusion-based Semantic Deflection
di: Liu, Qingyu, et al.
Pubblicazione: (2026)
di: Liu, Qingyu, et al.
Pubblicazione: (2026)
Discovering Command and Control (C2) Channels on Tor and Public Networks Using Reinforcement Learning
di: Wang, Cheng, et al.
Pubblicazione: (2024)
di: Wang, Cheng, et al.
Pubblicazione: (2024)
When Efficiency Backfires: Cascading LLMs Trigger Cascade Failure under Adversarial Attack
di: Sun, Zehan, et al.
Pubblicazione: (2026)
di: Sun, Zehan, et al.
Pubblicazione: (2026)
False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models
di: Jiang, Weipeng, et al.
Pubblicazione: (2026)
di: Jiang, Weipeng, et al.
Pubblicazione: (2026)
RESTRAIN: Reinforcement Learning-Based Secure Framework for Trigger-Action IoT Environment
di: Alam, Md Morshed, et al.
Pubblicazione: (2025)
di: Alam, Md Morshed, et al.
Pubblicazione: (2025)
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
di: Pang, Yan, et al.
Pubblicazione: (2025)
di: Pang, Yan, et al.
Pubblicazione: (2025)
JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
di: Chu, Junjie, et al.
Pubblicazione: (2025)
di: Chu, Junjie, et al.
Pubblicazione: (2025)
A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense
di: Zhai, Keke
Pubblicazione: (2024)
di: Zhai, Keke
Pubblicazione: (2024)
Documenti analoghi
-
Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood Discrepancy
di: Zhai, Shengfang, et al.
Pubblicazione: (2024) -
Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation Variation
di: Zhai, Shengfang, et al.
Pubblicazione: (2025) -
Purify Once, Edit Freely: Breaking Image Protections under Model Mismatch
di: Zhao, Qichen, et al.
Pubblicazione: (2026) -
Life-Cycle Routing Vulnerabilities of LLM Router
di: Lin, Qiqi, et al.
Pubblicazione: (2025) -
Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
di: Wang, Yuhao, et al.
Pubblicazione: (2025)