Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bonagiri, Akash, Li, Lucen, Oak, Rajvardhan, Babar, Zeerak, Wojcieszak, Magdalena, Chhabra, Anshuman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms
by: Oak, Rajvardhan, et al.
Published: (2025)
by: Oak, Rajvardhan, et al.
Published: (2025)
"Whose Side Are You On?" Estimating Ideology of Political and News Content Using Large Language Models and Few-shot Demonstration Selection
by: Haroon, Muhammad, et al.
Published: (2025)
by: Haroon, Muhammad, et al.
Published: (2025)
Incentivizing News Consumption on Social Media Platforms Using Large Language Models and Realistic Bot Accounts
by: Askari, Hadi, et al.
Published: (2024)
by: Askari, Hadi, et al.
Published: (2024)
"Hello, is this Anna?": Unpacking the Lifecycle of Pig-Butchering Scams
by: Oak, Rajvardhan, et al.
Published: (2025)
by: Oak, Rajvardhan, et al.
Published: (2025)
Understanding Underground Incentivized Review Services
by: Oak, Rajvardhan, et al.
Published: (2021)
by: Oak, Rajvardhan, et al.
Published: (2021)
MetaHarm: Harmful YouTube Video Dataset Annotated by Domain Experts, GPT-4-Turbo, and Crowdworkers
by: Jo, Wonjeong, et al.
Published: (2025)
by: Jo, Wonjeong, et al.
Published: (2025)
Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
by: Jo, Claire Wonjeong, et al.
Published: (2024)
by: Jo, Claire Wonjeong, et al.
Published: (2024)
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers
by: Achara, Akshit, et al.
Published: (2025)
by: Achara, Akshit, et al.
Published: (2025)
Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias
by: Chhabra, Anshuman, et al.
Published: (2024)
by: Chhabra, Anshuman, et al.
Published: (2024)
A Capabilities Approach to Studying Bias and Harm in Language Technologies
by: Nigatu, Hellina Hailu, et al.
Published: (2024)
by: Nigatu, Hellina Hailu, et al.
Published: (2024)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
by: Mendu, Sai Krishna, et al.
Published: (2025)
by: Mendu, Sai Krishna, et al.
Published: (2025)
Content Moderation Justice and Fairness on Social Media: Comparisons Across Different Contexts and Platforms
by: Cai, Jie, et al.
Published: (2024)
by: Cai, Jie, et al.
Published: (2024)
StopHC: A Harmful Content Detection and Mitigation Architecture for Social Media Platforms
by: Truică, Ciprian-Octavian, et al.
Published: (2024)
by: Truică, Ciprian-Octavian, et al.
Published: (2024)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation
by: Vitel, Dmytro, et al.
Published: (2025)
by: Vitel, Dmytro, et al.
Published: (2025)
The Unappreciated Role of Intent in Algorithmic Moderation of Social Media Content
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Wisdom of the LLM Crowd: A Large Scale Benchmark of Multi-Label U.S. Election-Related Harmful Social Media Content
by: Wang, Qile, et al.
Published: (2026)
by: Wang, Qile, et al.
Published: (2026)
Protecting Young Users on Social Media: Evaluating the Effectiveness of Content Moderation and Legal Safeguards on Video Sharing Platforms
by: Eltaher, Fatmaelzahraa, et al.
Published: (2025)
by: Eltaher, Fatmaelzahraa, et al.
Published: (2025)
Classist Tools: Social Class Correlates with Performance in NLP
by: Curry, Amanda Cercas, et al.
Published: (2024)
by: Curry, Amanda Cercas, et al.
Published: (2024)
Toward Accountable AI-Generated Content on Social Platforms: Steganographic Attribution and Multimodal Harm Detection
by: Guan, Xinlei, et al.
Published: (2026)
by: Guan, Xinlei, et al.
Published: (2026)
Personalizing Content Moderation on Social Media: User Perspectives on Moderation Choices, Interface Design, and Labor
by: Jhaver, Shagun, et al.
Published: (2023)
by: Jhaver, Shagun, et al.
Published: (2023)
From Anger to Joy: How Nationality Personas Shape Emotion Attribution in Large Language Models
by: Kamruzzaman, Mahammed, et al.
Published: (2025)
by: Kamruzzaman, Mahammed, et al.
Published: (2025)
Critical Challenges in Content Moderation for People Who Use Drugs (PWUD): Insights into Online Harm Reduction Practices from Moderators
by: Wang, Kaixuan, et al.
Published: (2025)
by: Wang, Kaixuan, et al.
Published: (2025)
Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
by: Mamidala, Rushitha Santhoshi, et al.
Published: (2025)
by: Mamidala, Rushitha Santhoshi, et al.
Published: (2025)
FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data
by: Abdelkadir, Nuredin Ali, et al.
Published: (2026)
by: Abdelkadir, Nuredin Ali, et al.
Published: (2026)
Governance of AI-Generated Content: A Case Study on Social Media Platforms
by: Gao, Lan, et al.
Published: (2026)
by: Gao, Lan, et al.
Published: (2026)
Understanding the Perceptions of Trigger Warning and Content Warning on Social Media Platforms in the U.S
by: Zhang, Xinyi, et al.
Published: (2025)
by: Zhang, Xinyi, et al.
Published: (2025)
Assessing LLMs for Zero-shot Abstractive Summarization Through the Lens of Relevance Paraphrasing
by: Askari, Hadi, et al.
Published: (2024)
by: Askari, Hadi, et al.
Published: (2024)
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening
by: Nahin, Shahriar Kabir, et al.
Published: (2026)
by: Nahin, Shahriar Kabir, et al.
Published: (2026)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
by: Gupta, Akash, et al.
Published: (2025)
by: Gupta, Akash, et al.
Published: (2025)
Closing the Motivation Gap: Incentives Enhance Visual Misinformation Discernment and Verification
by: Qian, Sijia, et al.
Published: (2026)
by: Qian, Sijia, et al.
Published: (2026)
Impoverished Language Technology: The Lack of (Social) Class in NLP
by: Curry, Amanda Cercas, et al.
Published: (2024)
by: Curry, Amanda Cercas, et al.
Published: (2024)
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
by: Monjur, Ocean, et al.
Published: (2026)
by: Monjur, Ocean, et al.
Published: (2026)
Longitudinal Monitoring of LLM Content Moderation of Social Issues
by: Dai, Yunlang, et al.
Published: (2025)
by: Dai, Yunlang, et al.
Published: (2025)
LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
by: Askari, Hadi, et al.
Published: (2025)
by: Askari, Hadi, et al.
Published: (2025)
Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
by: Gajewska, Ewelina, et al.
Published: (2026)
by: Gajewska, Ewelina, et al.
Published: (2026)
NLP4Gov: A Comprehensive Library for Computational Policy Analysis
by: Chakraborti, Mahasweta, et al.
Published: (2024)
by: Chakraborti, Mahasweta, et al.
Published: (2024)
Personal Moderation Configurations on Facebook: Exploring the Role of FoMO, Social Media Addiction, Norms, and Platform Trust
by: Jhaver, Shagun
Published: (2024)
by: Jhaver, Shagun
Published: (2024)
Polarized Online Discourse on Abortion: Frames and Hostile Expressions among Liberals and Conservatives
by: Rao, Ashwin, et al.
Published: (2023)
by: Rao, Ashwin, et al.
Published: (2023)
Similar Items
-
Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms
by: Oak, Rajvardhan, et al.
Published: (2025) -
"Whose Side Are You On?" Estimating Ideology of Political and News Content Using Large Language Models and Few-shot Demonstration Selection
by: Haroon, Muhammad, et al.
Published: (2025) -
Incentivizing News Consumption on Social Media Platforms Using Large Language Models and Realistic Bot Accounts
by: Askari, Hadi, et al.
Published: (2024) -
"Hello, is this Anna?": Unpacking the Lifecycle of Pig-Butchering Scams
by: Oak, Rajvardhan, et al.
Published: (2025) -
Understanding Underground Incentivized Review Services
by: Oak, Rajvardhan, et al.
Published: (2021)