Advancing Content Moderation: Evaluating Large Language Models for Detecting Sensitive Content Across Text, Images, and Videos
Fuente:
arXiv
Saved in:
| Main Authors: | AlDahoul, Nouar, Tan, Myles Joshua Toledo, Kasireddy, Harishwar Reddy, Zaki, Yasir |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Vision Language Models for Facial Attribute Recognition: Emotion, Race, Gender, and Age
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
A Novel BERT-based Classifier to Detect Political Leaning of YouTube Videos based on their Titles
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
AI-generated faces influence gender stereotypes and racial homogenization
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
Advancing Vehicle Plate Recognition: Multitasking Visual Language Models with VehiclePaliGemma
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
Self-Reflection Makes Large Language Models Safer, Less Biased, and Ideologically Neutral
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
A Conceptual Exploration of Generative AI-Induced Cognitive Dissonance and its Emergence in University-Level Academic Writing
by: Seran, Carl Errol, et al.
Published: (2025)
by: Seran, Carl Errol, et al.
Published: (2025)
Inclusive content reduces racial and gender biases, yet non-inclusive content dominates popular culture
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
by: Jan, Essa, et al.
Published: (2024)
by: Jan, Essa, et al.
Published: (2024)
Fine-tuned Vision Language Model for Localization of Parasitic Eggs in Microscopic Images
by: Sien, Chan Hao, et al.
Published: (2026)
by: Sien, Chan Hao, et al.
Published: (2026)
A Longitudinal Analysis of Racial and Gender Bias in New York Times and Fox News Images and Articles
by: Ibrahim, Hazem, et al.
Published: (2024)
by: Ibrahim, Hazem, et al.
Published: (2024)
Empowering the Grid: Collaborative Edge Artificial Intelligence for Decentralized Energy Systems
by: Paula Jr, Eddie de, et al.
Published: (2025)
by: Paula Jr, Eddie de, et al.
Published: (2025)
Real-Time Human Detection for Aerial Captured Video Sequences via Deep Models
by: AlDahoul, Nouar, et al.
Published: (2026)
by: AlDahoul, Nouar, et al.
Published: (2026)
Neutralizing the Narrative: AI-Powered Debiasing of Online News Articles
by: Kuo, Chen Wei, et al.
Published: (2025)
by: Kuo, Chen Wei, et al.
Published: (2025)
Enhancing Password Security Through a High-Accuracy Scoring Framework Using Random Forests
by: Mazelan, Muhammed El Mustaqeem, et al.
Published: (2025)
by: Mazelan, Muhammed El Mustaqeem, et al.
Published: (2025)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
by: Aldahoul, Nouar, et al.
Published: (2025)
by: Aldahoul, Nouar, et al.
Published: (2025)
Can Personalized Medicine Coexist with Health Equity? Examining the Cost Barrier and Ethical Implications
by: Francisco, Kishi Kobe Yee, et al.
Published: (2024)
by: Francisco, Kishi Kobe Yee, et al.
Published: (2024)
Semantic-Aware Advanced Persistent Threat Detection Using Autoencoders on LLM-Encoded System Logs
by: Mohammed, Waleed Khan, et al.
Published: (2026)
by: Mohammed, Waleed Khan, et al.
Published: (2026)
KidneyHFM-Eval: Validation of Histopathology Foundation Models
by: Kasireddy, Harishwar Reddy, et al.
Published: (2026)
by: Kasireddy, Harishwar Reddy, et al.
Published: (2026)
Metamorphic Testing for Fairness Evaluation in Large Language Models: Identifying Intersectional Bias in LLaMA and GPT
by: Reddy, Harishwar, et al.
Published: (2025)
by: Reddy, Harishwar, et al.
Published: (2025)
An explainable Recursive Feature Elimination to detect Advanced Persistent Threats using Random Forest classifier
by: Mutalib, Noor Hazlina Abdul, et al.
Published: (2025)
by: Mutalib, Noor Hazlina Abdul, et al.
Published: (2025)
Who Gets Seen in the Age of AI? Adoption Patterns of Large Language Models in Scholarly Writing and Citation Outcomes
by: Khan, Farhan Kamrul, et al.
Published: (2025)
by: Khan, Farhan Kamrul, et al.
Published: (2025)
Large Language Models are often politically extreme, usually ideologically inconsistent, and persuasive even in informational contexts
by: Aldahoul, Nouar, et al.
Published: (2025)
by: Aldahoul, Nouar, et al.
Published: (2025)
Exploring the Boundaries of Content Moderation in Text-to-Image Generation
by: Riccio, Piera, et al.
Published: (2024)
by: Riccio, Piera, et al.
Published: (2024)
Scaling Reinforcement Learning for Content Moderation with Large Language Models
by: Firooz, Hamed, et al.
Published: (2025)
by: Firooz, Hamed, et al.
Published: (2025)
Watch Your Language: Investigating Content Moderation with Large Language Models
by: Kumar, Deepak, et al.
Published: (2023)
by: Kumar, Deepak, et al.
Published: (2023)
Evaluating the Efficacy of Next.js: A Comparative Analysis with React.js on Performance, SEO, and Global Network Equity
by: Pati, Swostik, et al.
Published: (2025)
by: Pati, Swostik, et al.
Published: (2025)
Addressing Intersectionality, Explainability, and Ethics in AI-Driven Diagnostics: A Rebuttal and Call for Transdiciplinary Action
by: Tan, Myles Joshua Toledo, et al.
Published: (2025)
by: Tan, Myles Joshua Toledo, et al.
Published: (2025)
Quantum Bit Threads and the Entropohedron
by: Headrick, Matthew, et al.
Published: (2025)
by: Headrick, Matthew, et al.
Published: (2025)
Policy-as-Prompt: Rethinking Content Moderation in the Age of Large Language Models
by: Palla, Konstantina, et al.
Published: (2025)
by: Palla, Konstantina, et al.
Published: (2025)
Legilimens: Practical and Unified Content Moderation for Large Language Model Services
by: Wu, Jialin, et al.
Published: (2024)
by: Wu, Jialin, et al.
Published: (2024)
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
by: Dinuta, Eduard Stefan, et al.
Published: (2025)
by: Dinuta, Eduard Stefan, et al.
Published: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
Content Moderation Futures
by: Blackwell, Lindsay
Published: (2025)
by: Blackwell, Lindsay
Published: (2025)
Shaping Integrity: Why Generative Artificial Intelligence Does Not Have to Undermine Education
by: Tan, Myles Joshua Toledo, et al.
Published: (2024)
by: Tan, Myles Joshua Toledo, et al.
Published: (2024)
On Demographic Transformation: Why We Need to Think Beyond Silos
by: Maravilla, Nicholle Mae Amor Tan, et al.
Published: (2025)
by: Maravilla, Nicholle Mae Amor Tan, et al.
Published: (2025)
From Text to Source: Results in Detecting Large Language Model-Generated Content
by: Antoun, Wissam, et al.
Published: (2023)
by: Antoun, Wissam, et al.
Published: (2023)
PixLift: Accelerating Web Browsing via AI Upscaling
by: Atinafu, Yonas, et al.
Published: (2025)
by: Atinafu, Yonas, et al.
Published: (2025)
Similar Items
-
Exploring Vision Language Models for Facial Attribute Recognition: Emotion, Race, Gender, and Age
by: AlDahoul, Nouar, et al.
Published: (2024) -
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
by: AlDahoul, Nouar, et al.
Published: (2025) -
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
by: AlDahoul, Nouar, et al.
Published: (2025) -
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
by: AlDahoul, Nouar, et al.
Published: (2025) -
A Novel BERT-based Classifier to Detect Political Leaning of YouTube Videos based on their Titles
by: AlDahoul, Nouar, et al.
Published: (2024)