UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images
Fuente:
arXiv
Saved in:
| Main Authors: | Qu, Yiting, Shen, Xinyue, Wu, Yixin, Backes, Michael, Zannettou, Savvas, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
by: Shen, Xinyue, et al.
Published: (2025)
by: Shen, Xinyue, et al.
Published: (2025)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
by: Wu, Yixin, et al.
Published: (2023)
by: Wu, Yixin, et al.
Published: (2023)
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
by: Qu, Yiting, et al.
Published: (2025)
by: Qu, Yiting, et al.
Published: (2025)
Prompt Stealing Attacks Against Text-to-Image Generation Models
by: Shen, Xinyue, et al.
Published: (2023)
by: Shen, Xinyue, et al.
Published: (2023)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
by: Qu, Yiting, et al.
Published: (2025)
by: Qu, Yiting, et al.
Published: (2025)
Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media
by: Sun, Zhen, et al.
Published: (2024)
by: Sun, Zhen, et al.
Published: (2024)
"Humans welcome to observe": A First Look at the Agent Social Network Moltbook
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
by: Wu, Yixin, et al.
Published: (2024)
by: Wu, Yixin, et al.
Published: (2024)
Auditing the Compliance and Enforcement of Twitter's Advertising Policy
by: Vekaria, Yash, et al.
Published: (2023)
by: Vekaria, Yash, et al.
Published: (2023)
From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection
by: Liang, Mengfei, et al.
Published: (2025)
by: Liang, Mengfei, et al.
Published: (2025)
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
Voice Jailbreak Attacks Against GPT-4o
by: Shen, Xinyue, et al.
Published: (2024)
by: Shen, Xinyue, et al.
Published: (2024)
AI-Generated Faces in the Real World: A Large-Scale Case Study of Twitter Profile Images
by: Ricker, Jonas, et al.
Published: (2024)
by: Ricker, Jonas, et al.
Published: (2024)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Privacy Intelligence: A Survey on Image Privacy in Online Social Networks
by: Liu, Chi, et al.
Published: (2020)
by: Liu, Chi, et al.
Published: (2020)
MGTBench: Benchmarking Machine-Generated Text Detection
by: He, Xinlei, et al.
Published: (2023)
by: He, Xinlei, et al.
Published: (2023)
Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation
by: Keita, Mamadou, et al.
Published: (2026)
by: Keita, Mamadou, et al.
Published: (2026)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
$OIDC^2$: Open Identity Certification with OpenID Connect
by: Primbs, Jonas, et al.
Published: (2023)
by: Primbs, Jonas, et al.
Published: (2023)
Blockchain Takeovers in Web 3.0: An Empirical Study on the TRON-Steem Incident
by: Li, Chao, et al.
Published: (2024)
by: Li, Chao, et al.
Published: (2024)
UniDetect: LLM-Driven Universal Fraud Detection across Heterogeneous Blockchains
by: Miao, Shuyi, et al.
Published: (2026)
by: Miao, Shuyi, et al.
Published: (2026)
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
by: Shen, Xinyue, et al.
Published: (2025)
by: Shen, Xinyue, et al.
Published: (2025)
Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
by: Chen, Zeyuan, et al.
Published: (2026)
by: Chen, Zeyuan, et al.
Published: (2026)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
by: Meng, Xiangtao, et al.
Published: (2025)
by: Meng, Xiangtao, et al.
Published: (2025)
"That's another doom I haven't thought about": A User Study on AI Labels as a Safeguard Against Image-Based Misinformation
by: Höltervennhoff, Sandra, et al.
Published: (2025)
by: Höltervennhoff, Sandra, et al.
Published: (2025)
Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications
by: Wu, Yixin, et al.
Published: (2025)
by: Wu, Yixin, et al.
Published: (2025)
Peering Behind the Shield: Guardrail Identification in Large Language Models
by: Yang, Ziqing, et al.
Published: (2025)
by: Yang, Ziqing, et al.
Published: (2025)
MoltGraph: A Longitudinal Temporal Graph Dataset of Moltbook for Coordinated-Agent Detection
by: Mukherjee, Kunal, et al.
Published: (2026)
by: Mukherjee, Kunal, et al.
Published: (2026)
Resource Allocation and Secure Wireless Communication in the Large Model-based Mobile Edge Computing System
by: Wang, Zefan, et al.
Published: (2024)
by: Wang, Zefan, et al.
Published: (2024)
SoK: Analysis of Privacy Risks and Mitigation in Online Propaganda Detection through the PROMPT Framework
by: Goswami, Dhiman, et al.
Published: (2026)
by: Goswami, Dhiman, et al.
Published: (2026)
Empirical Network Structure of Malicious Programs
by: Musgrave, John, et al.
Published: (2022)
by: Musgrave, John, et al.
Published: (2022)
Hiding from Facebook: An Encryption Protocol resistant to Correlation Attacks
by: Liu, Chen-Da, et al.
Published: (2024)
by: Liu, Chen-Da, et al.
Published: (2024)
Detecting and Understanding the Promotion of Illicit Goods and Services on Twitter
by: Wang, Hongyu, et al.
Published: (2024)
by: Wang, Hongyu, et al.
Published: (2024)
Making Privacy-preserving Federated Graph Analytics with Strong Guarantees Practical (for Certain Queries)
by: Liu, Kunlong, et al.
Published: (2024)
by: Liu, Kunlong, et al.
Published: (2024)
Unfair Mistakes on Social Media: How Demographic Characteristics influence Authorship Attribution
by: Wyss, Jasmin, et al.
Published: (2025)
by: Wyss, Jasmin, et al.
Published: (2025)
Username Squatting on Online Social Networks: A Study on X
by: Lepipas, Anastasios, et al.
Published: (2024)
by: Lepipas, Anastasios, et al.
Published: (2024)
Cross-border Exchange of CBDCs using Layer-2 Blockchain
by: Gogol, Krzysztof, et al.
Published: (2023)
by: Gogol, Krzysztof, et al.
Published: (2023)
Integrating Network and Attack Graphs for Service-Centric Impact Analysis
by: Herttuainen, Joni, et al.
Published: (2025)
by: Herttuainen, Joni, et al.
Published: (2025)
Similar Items
-
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
by: Shen, Xinyue, et al.
Published: (2025) -
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
by: Wu, Yixin, et al.
Published: (2023) -
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
by: Qu, Yiting, et al.
Published: (2025) -
Prompt Stealing Attacks Against Text-to-Image Generation Models
by: Shen, Xinyue, et al.
Published: (2023) -
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
by: Chu, Junjie, et al.
Published: (2026)