Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
Fuente:
arXiv
Saved in:
| Main Authors: | Qu, Yiting, Yang, Ziqing, Ma, Yihan, Backes, Michael, Zannettou, Savvas, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
by: Shen, Xinyue, et al.
Published: (2025)
by: Shen, Xinyue, et al.
Published: (2025)
UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images
by: Qu, Yiting, et al.
Published: (2024)
by: Qu, Yiting, et al.
Published: (2024)
From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection
by: Liang, Mengfei, et al.
Published: (2025)
by: Liang, Mengfei, et al.
Published: (2025)
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
by: Yang, Ziqing, et al.
Published: (2026)
by: Yang, Ziqing, et al.
Published: (2026)
Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face Detection
by: Zhu, Delong, et al.
Published: (2024)
by: Zhu, Delong, et al.
Published: (2024)
When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigm
by: Leng, Ye, et al.
Published: (2026)
by: Leng, Ye, et al.
Published: (2026)
Still Camouflage, Moving Illusion: View-Induced Trajectory Manipulation in Autonomous Driving
by: Ju, Shuo, et al.
Published: (2026)
by: Ju, Shuo, et al.
Published: (2026)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Delving into Decision-based Black-box Attacks on Semantic Segmentation
by: Chen, Zhaoyu, et al.
Published: (2024)
by: Chen, Zhaoyu, et al.
Published: (2024)
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
DivTrackee versus DynTracker: Promoting Diversity in Anti-Facial Recognition against Dynamic FR Strategy
by: Fan, Wenshu, et al.
Published: (2025)
by: Fan, Wenshu, et al.
Published: (2025)
A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Hiding-in-Plain-Sight (HiPS) Attack on CLIP for Targetted Object Removal from Images
by: Daw, Arka, et al.
Published: (2024)
by: Daw, Arka, et al.
Published: (2024)
Detecting Malicious Concepts without Image Generation in AI-Generated Content (AIGC)
by: Xu, Kun, et al.
Published: (2025)
by: Xu, Kun, et al.
Published: (2025)
FreqCross: A Multi-Modal Frequency-Spatial Fusion Network for Robust Detection of Stable Diffusion 3.5 Generated Images
by: Yang, Guang
Published: (2025)
by: Yang, Guang
Published: (2025)
LiteUpdate: A Lightweight Framework for Updating AI-Generated Image Detectors
by: Lu, Jiajie, et al.
Published: (2025)
by: Lu, Jiajie, et al.
Published: (2025)
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models
by: Liu, Yule, et al.
Published: (2026)
by: Liu, Yule, et al.
Published: (2026)
DeepSight: An All-in-One LM Safety Toolkit
by: Zhang, Bo, et al.
Published: (2026)
by: Zhang, Bo, et al.
Published: (2026)
Anti-Tamper Protection for Unauthorized Individual Image Generation
by: Li, Zelin, et al.
Published: (2025)
by: Li, Zelin, et al.
Published: (2025)
The Devil's Advocate: Shattering the Illusion of Unexploitable Data using Diffusion Models
by: Dolatabadi, Hadi M., et al.
Published: (2023)
by: Dolatabadi, Hadi M., et al.
Published: (2023)
SKeDA: A Generative Watermarking Framework for Text-to-video Diffusion Models
by: Yang, Yang, et al.
Published: (2026)
by: Yang, Yang, et al.
Published: (2026)
Vulnerabilities in AI-generated Image Detection: The Challenge of Adversarial Attacks
by: Diao, Yunfeng, et al.
Published: (2024)
by: Diao, Yunfeng, et al.
Published: (2024)
AI-Generated Video Detection via Spatio-Temporal Anomaly Learning
by: Bai, Jianfa, et al.
Published: (2024)
by: Bai, Jianfa, et al.
Published: (2024)
Generative AI-Based Effective Malware Detection for Embedded Computing Systems
by: Kasarapu, Sreenitha, et al.
Published: (2024)
by: Kasarapu, Sreenitha, et al.
Published: (2024)
MIRROR: Manifold Ideal Reference ReconstructOR for Generalizable AI-Generated Image Detection
by: Liu, Ruiqi, et al.
Published: (2026)
by: Liu, Ruiqi, et al.
Published: (2026)
Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features
by: Wu, Shangbo, et al.
Published: (2025)
by: Wu, Shangbo, et al.
Published: (2025)
CLIP-Flow: A Universal Discriminator for AI-Generated Images Inspired by Anomaly Detection
by: Yuan, Zhipeng, et al.
Published: (2025)
by: Yuan, Zhipeng, et al.
Published: (2025)
SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts
by: Wu, Xiaodong, et al.
Published: (2025)
by: Wu, Xiaodong, et al.
Published: (2025)
Transferable Dual-Domain Feature Importance Attack against AI-Generated Image Detector
by: Zhu, Weiheng, et al.
Published: (2025)
by: Zhu, Weiheng, et al.
Published: (2025)
The Orthogonal Vulnerabilities of Generative AI Watermarks: A Comparative Empirical Benchmark of Spatial and Latent Provenance
by: Yu, Jesse, et al.
Published: (2026)
by: Yu, Jesse, et al.
Published: (2026)
DeeCLIP: A Robust and Generalizable Transformer-Based Framework for Detecting AI-Generated Images
by: Keita, Mamadou, et al.
Published: (2025)
by: Keita, Mamadou, et al.
Published: (2025)
Color Matters: Demosaicing-Guided Color Correlation Training for Generalizable AI-Generated Image Detection
by: Zhong, Nan, et al.
Published: (2026)
by: Zhong, Nan, et al.
Published: (2026)
TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models
by: Zhang, Chaoshuo, et al.
Published: (2026)
by: Zhang, Chaoshuo, et al.
Published: (2026)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
Quantifying the Risk of Transferred Black Box Attacks
by: Cox, Disesdi Susanna, et al.
Published: (2025)
by: Cox, Disesdi Susanna, et al.
Published: (2025)
Exploit the Leak: Understanding Risks in Biometric Matchers
by: Durbet, Axel, et al.
Published: (2023)
by: Durbet, Axel, et al.
Published: (2023)
Membership Inference Attack Against Masked Image Modeling
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey
by: Cao, Jie, et al.
Published: (2025)
by: Cao, Jie, et al.
Published: (2025)
Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models
by: Zhang, Rui, et al.
Published: (2025)
by: Zhang, Rui, et al.
Published: (2025)
3D-ANC: Adaptive Neural Collapse for Robust 3D Point Cloud Recognition
by: Huang, Yuanmin, et al.
Published: (2025)
by: Huang, Yuanmin, et al.
Published: (2025)
Similar Items
-
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
by: Shen, Xinyue, et al.
Published: (2025) -
UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images
by: Qu, Yiting, et al.
Published: (2024) -
From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection
by: Liang, Mengfei, et al.
Published: (2025) -
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
by: Yang, Ziqing, et al.
Published: (2026) -
Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face Detection
by: Zhu, Delong, et al.
Published: (2024)