Moderating Illicit Online Image Promotion for Unsafe User-Generated Content Games Using Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Keyan, Utkarsh, Ayush, Ding, Wenbo, Ondracek, Isabelle, Zhao, Ziming, Freeman, Guo, Vishwamitra, Nishant, Hu, Hongxin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Moderating New Waves of Online Hate with Chain-of-Thought Reasoning in Large Language Models
by: Vishwamitra, Nishant, et al.
Published: (2023)
by: Vishwamitra, Nishant, et al.
Published: (2023)
Image Safeguarding: Reasoning with Conditional Vision Language Model and Obfuscating Unsafe Content Counterfactually
by: Bethany, Mazal, et al.
Published: (2024)
by: Bethany, Mazal, et al.
Published: (2024)
Moderating Embodied Cyber Threats Using Generative AI
by: Guo, Keyan, et al.
Published: (2024)
by: Guo, Keyan, et al.
Published: (2024)
An Investigation of Large Language Models for Real-World Hate Speech Detection
by: Guo, Keyan, et al.
Published: (2024)
by: Guo, Keyan, et al.
Published: (2024)
AI-Cybersecurity Education Through Designing AI-based Cyberharassment Detection Lab
by: Okpala, Ebuka, et al.
Published: (2024)
by: Okpala, Ebuka, et al.
Published: (2024)
Detection of Illicit Content on Online Marketplaces using Large Language Models
by: Tran, Quoc Khoa, et al.
Published: (2026)
by: Tran, Quoc Khoa, et al.
Published: (2026)
Can Reinforcement Learning Unlock the Hidden Dangers in Aligned Large Language Models?
by: Karkevandi, Mohammad Bahrami, et al.
Published: (2024)
by: Karkevandi, Mohammad Bahrami, et al.
Published: (2024)
Users Views about the Usability of Digital Libraries
by: Koohang, Alex, et al.
Published: (2005)
by: Koohang, Alex, et al.
Published: (2005)
Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon Detection
by: Song, Minkyoo, et al.
Published: (2025)
by: Song, Minkyoo, et al.
Published: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
Promoting Online Safety by Simulating Unsafe Conversations with LLMs
by: Hoffman, Owen, et al.
Published: (2025)
by: Hoffman, Owen, et al.
Published: (2025)
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
by: Foo, Jessica, et al.
Published: (2024)
by: Foo, Jessica, et al.
Published: (2024)
CAMOUFLAGE: Exploiting Misinformation Detection Systems Through LLM-driven Adversarial Claim Transformation
by: Bethany, Mazal, et al.
Published: (2025)
by: Bethany, Mazal, et al.
Published: (2025)
Deciphering Textual Authenticity: A Generalized Strategy through the Lens of Large Language Semantics for Detecting Human vs. Machine-Generated Text
by: Bethany, Mazal, et al.
Published: (2024)
by: Bethany, Mazal, et al.
Published: (2024)
Hidden in Plain Sight: Detecting Illicit Massage Businesses from Mobility Data
by: Shomali, Roya, et al.
Published: (2026)
by: Shomali, Roya, et al.
Published: (2026)
Lateral Phishing With Large Language Models: A Large Organization Comparative Study
by: Bethany, Mazal, et al.
Published: (2024)
by: Bethany, Mazal, et al.
Published: (2024)
Improving Regulatory Oversight in Online Content Moderation
by: Tessa, Benedetta, et al.
Published: (2025)
by: Tessa, Benedetta, et al.
Published: (2025)
Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models
by: Cui, Chenhang, et al.
Published: (2024)
by: Cui, Chenhang, et al.
Published: (2024)
Enhancing Event Reasoning in Large Language Models through Instruction Fine-Tuning with Semantic Causal Graphs
by: Bethany, Mazal, et al.
Published: (2024)
by: Bethany, Mazal, et al.
Published: (2024)
Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents
by: Talokar, Nivya, et al.
Published: (2026)
by: Talokar, Nivya, et al.
Published: (2026)
Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation
by: Samory, Mattia, et al.
Published: (2025)
by: Samory, Mattia, et al.
Published: (2025)
Image Recognition with Online Lightweight Vision Transformer: A Survey
by: Zhang, Zherui, et al.
Published: (2025)
by: Zhang, Zherui, et al.
Published: (2025)
Impact of Stricter Content Moderation on Parler's Users' Discourse
by: Kumarswamy, Nihal, et al.
Published: (2023)
by: Kumarswamy, Nihal, et al.
Published: (2023)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Personalizing Content Moderation on Social Media: User Perspectives on Moderation Choices, Interface Design, and Labor
by: Jhaver, Shagun, et al.
Published: (2023)
by: Jhaver, Shagun, et al.
Published: (2023)
Content Moderation Futures
by: Blackwell, Lindsay
Published: (2025)
by: Blackwell, Lindsay
Published: (2025)
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
by: Ghosh, Shaona, et al.
Published: (2024)
by: Ghosh, Shaona, et al.
Published: (2024)
Toward Early Quality Assessment of Text-to-Image Diffusion Models
by: Guo, Huanlei, et al.
Published: (2026)
by: Guo, Huanlei, et al.
Published: (2026)
LLM Content Moderation and User Satisfaction: Evidence from Response Refusals in Chatbot Arena
by: Pasch, Stefan
Published: (2025)
by: Pasch, Stefan
Published: (2025)
Critical Challenges in Content Moderation for People Who Use Drugs (PWUD): Insights into Online Harm Reduction Practices from Moderators
by: Wang, Kaixuan, et al.
Published: (2025)
by: Wang, Kaixuan, et al.
Published: (2025)
Exploring the Boundaries of Content Moderation in Text-to-Image Generation
by: Riccio, Piera, et al.
Published: (2024)
by: Riccio, Piera, et al.
Published: (2024)
Quantifying Feature Importance for Online Content Moderation
by: Tessa, Benedetta, et al.
Published: (2025)
by: Tessa, Benedetta, et al.
Published: (2025)
Illicit object detection in X-ray images using Vision Transformers
by: Cani, Jorgen, et al.
Published: (2024)
by: Cani, Jorgen, et al.
Published: (2024)
Participation Incentives in Online Cooperative Games
by: Aziz, Haris, et al.
Published: (2025)
by: Aziz, Haris, et al.
Published: (2025)
Unsafe2Safe: Controllable Image Anonymization for Downstream Utility
by: Dinh, Mih, et al.
Published: (2026)
by: Dinh, Mih, et al.
Published: (2026)
Optimal Pseudorandom Generators for Low-Degree Polynomials Over Moderately Large Fields
by: Dwivedi, Ashish, et al.
Published: (2024)
by: Dwivedi, Ashish, et al.
Published: (2024)
Throwaway Accounts and Moderation on Reddit
by: Guo, Cheng, et al.
Published: (2025)
by: Guo, Cheng, et al.
Published: (2025)
Attention Shift: Steering AI Away from Unsafe Content
by: Garg, Shivank, et al.
Published: (2024)
by: Garg, Shivank, et al.
Published: (2024)
ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation
by: Wu, Mengyang, et al.
Published: (2024)
by: Wu, Mengyang, et al.
Published: (2024)
Let Community Rules Be Reflected in Online Content Moderation
by: Xin, Wangjiaxuan, et al.
Published: (2024)
by: Xin, Wangjiaxuan, et al.
Published: (2024)
Similar Items
-
Moderating New Waves of Online Hate with Chain-of-Thought Reasoning in Large Language Models
by: Vishwamitra, Nishant, et al.
Published: (2023) -
Image Safeguarding: Reasoning with Conditional Vision Language Model and Obfuscating Unsafe Content Counterfactually
by: Bethany, Mazal, et al.
Published: (2024) -
Moderating Embodied Cyber Threats Using Generative AI
by: Guo, Keyan, et al.
Published: (2024) -
An Investigation of Large Language Models for Real-World Hate Speech Detection
by: Guo, Keyan, et al.
Published: (2024) -
AI-Cybersecurity Education Through Designing AI-based Cyberharassment Detection Lab
by: Okpala, Ebuka, et al.
Published: (2024)