TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chaoshuo, Liang, Yibo, Tian, Mengke, Lin, Chenhao, Zhao, Zhengyu, Yang, Le, Zhang, Chong, Zhang, Yang, Shen, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Privacy on the Fly: A Predictive Adversarial Transformation Network for Mobile Sensor Data
by: Song, Tianle, et al.
Published: (2025)
by: Song, Tianle, et al.
Published: (2025)
Systematic Categorization, Construction and Evaluation of New Attacks against Multi-modal Mobile GUI Agents
by: Yang, Yulong, et al.
Published: (2024)
by: Yang, Yulong, et al.
Published: (2024)
The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
by: Liu, Mingrui, et al.
Published: (2025)
by: Liu, Mingrui, et al.
Published: (2025)
A Survey on Adversarial Machine Learning for Code Data: Realistic Threats, Countermeasures, and Interpretations
by: Yang, Yulong, et al.
Published: (2024)
by: Yang, Yulong, et al.
Published: (2024)
LESSON: Multi-Label Adversarial False Data Injection Attack for Deep Learning Locational Detection
by: Tian, Jiwei, et al.
Published: (2024)
by: Tian, Jiwei, et al.
Published: (2024)
A Comparative Study of Fuzzers and Static Analysis Tools for Finding Memory Unsafety in C and C++
by: Hassler, Keno, et al.
Published: (2025)
by: Hassler, Keno, et al.
Published: (2025)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
by: Wu, Yixin, et al.
Published: (2023)
by: Wu, Yixin, et al.
Published: (2023)
Quantization Aware Attack: Enhancing Transferable Adversarial Attacks by Model Quantization
by: Yang, Yulong, et al.
Published: (2023)
by: Yang, Yulong, et al.
Published: (2023)
MGTBench: Benchmarking Machine-Generated Text Detection
by: He, Xinlei, et al.
Published: (2023)
by: He, Xinlei, et al.
Published: (2023)
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023)
by: Huang, Hai, et al.
Published: (2023)
Awakening the Hydra: Stabilizing Multi-Concept Backdoor Injection in Text-to-Image Diffusion Models
by: Wang, Kai, et al.
Published: (2026)
by: Wang, Kai, et al.
Published: (2026)
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
by: Lin, Chenhao, et al.
Published: (2025)
by: Lin, Chenhao, et al.
Published: (2025)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
by: Wu, Yixin, et al.
Published: (2024)
by: Wu, Yixin, et al.
Published: (2024)
Typographic Attacks in a Multi-Image Setting
by: Wang, Xiaomeng, et al.
Published: (2025)
by: Wang, Xiaomeng, et al.
Published: (2025)
Prompt Stealing Attacks Against Text-to-Image Generation Models
by: Shen, Xinyue, et al.
Published: (2023)
by: Shen, Xinyue, et al.
Published: (2023)
Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
by: Zheng, Junhao, et al.
Published: (2025)
by: Zheng, Junhao, et al.
Published: (2025)
Robustness Over Time: Understanding Adversarial Examples' Effectiveness on Longitudinal Versions of Large Language Models
by: Liu, Yugeng, et al.
Published: (2023)
by: Liu, Yugeng, et al.
Published: (2023)
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
by: Zhu, Zhenhao, et al.
Published: (2026)
by: Zhu, Zhenhao, et al.
Published: (2026)
Traceable AI-driven Avatars Using Multi-factors of Physical World and Metaverse
by: Yang, Kedi, et al.
Published: (2024)
by: Yang, Kedi, et al.
Published: (2024)
Prediction Inconsistency Helps Achieve Generalizable Detection of Adversarial Examples
by: Han, Sicong, et al.
Published: (2025)
by: Han, Sicong, et al.
Published: (2025)
Improving Adversarial Transferability on Vision Transformers via Forward Propagation Refinement
by: Ren, Yuchen, et al.
Published: (2025)
by: Ren, Yuchen, et al.
Published: (2025)
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
by: Xu, Wenzhuo, et al.
Published: (2026)
by: Xu, Wenzhuo, et al.
Published: (2026)
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
by: Zhao, Shiqian, et al.
Published: (2025)
by: Zhao, Shiqian, et al.
Published: (2025)
JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
by: Chu, Junjie, et al.
Published: (2025)
by: Chu, Junjie, et al.
Published: (2025)
Generalizable Targeted Data Poisoning against Varying Physical Objects
by: Chen, Zhizhen, et al.
Published: (2024)
by: Chen, Zhizhen, et al.
Published: (2024)
Physical 3D Adversarial Attacks against Monocular Depth Estimation in Autonomous Driving
by: Zheng, Junhao, et al.
Published: (2024)
by: Zheng, Junhao, et al.
Published: (2024)
Espresso: Robust Concept Filtering in Text-to-Image Models
by: Das, Anudeep, et al.
Published: (2024)
by: Das, Anudeep, et al.
Published: (2024)
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation Variation
by: Zhai, Shengfang, et al.
Published: (2025)
by: Zhai, Shengfang, et al.
Published: (2025)
CoreUnlearn: Rethinking Concept Unlearning through Disentangled Component-Level Erasure in Text-guided Diffusion Models
by: Zhao, Mengnan, et al.
Published: (2026)
by: Zhao, Mengnan, et al.
Published: (2026)
Don't Trust Your Upstream: Exploiting LLM Multi-Agent System via Topology-Guided Adversarial Propagation
by: Liang, Ruichao, et al.
Published: (2025)
by: Liang, Ruichao, et al.
Published: (2025)
Combinational Backdoor Attack against Customized Text-to-Image Models
by: Jiang, Wenbo, et al.
Published: (2024)
by: Jiang, Wenbo, et al.
Published: (2024)
Improving Integrated Gradient-based Transferable Adversarial Examples by Refining the Integration Path
by: Ren, Yuchen, et al.
Published: (2024)
by: Ren, Yuchen, et al.
Published: (2024)
Provably Robust Multi-bit Watermarking for AI-generated Text
by: Qu, Wenjie, et al.
Published: (2024)
by: Qu, Wenjie, et al.
Published: (2024)
When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigm
by: Leng, Ye, et al.
Published: (2026)
by: Leng, Ye, et al.
Published: (2026)
Two-layer consensus based on master-slave consortium chain data sharing for Internet of Vehicles
by: Zhao, Feng, et al.
Published: (2024)
by: Zhao, Feng, et al.
Published: (2024)
Mitigating the Backdoor Effect for Multi-Task Model Merging via Safety-Aware Subspace
by: Yang, Jinluan, et al.
Published: (2024)
by: Yang, Jinluan, et al.
Published: (2024)
Concept Unlearning by Modeling Key Steps of Diffusion Process
by: Zhang, Chaoshuo, et al.
Published: (2025)
by: Zhang, Chaoshuo, et al.
Published: (2025)
Toward Web 4.0: Bidirectional Trust between AI Agents and Blockchain
by: Xia, Yunfeng, et al.
Published: (2026)
by: Xia, Yunfeng, et al.
Published: (2026)
Evaluating Concept Filtering Defenses against Child Sexual Abuse Material Generation by Text-to-Image Models
by: Cretu, Ana-Maria, et al.
Published: (2025)
by: Cretu, Ana-Maria, et al.
Published: (2025)
Similar Items
-
Privacy on the Fly: A Predictive Adversarial Transformation Network for Mobile Sensor Data
by: Song, Tianle, et al.
Published: (2025) -
Systematic Categorization, Construction and Evaluation of New Attacks against Multi-modal Mobile GUI Agents
by: Yang, Yulong, et al.
Published: (2024) -
The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
by: Liu, Mingrui, et al.
Published: (2025) -
A Survey on Adversarial Machine Learning for Code Data: Realistic Threats, Countermeasures, and Interpretations
by: Yang, Yulong, et al.
Published: (2024) -
LESSON: Multi-Label Adversarial False Data Injection Attack for Deep Learning Locational Detection
by: Tian, Jiwei, et al.
Published: (2024)