Watch Your Language: Investigating Content Moderation with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Deepak, AbuHashem, Yousef, Durumeric, Zakir
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909075412877312
author Kumar, Deepak
AbuHashem, Yousef
Durumeric, Zakir
author_facet Kumar, Deepak
AbuHashem, Yousef
Durumeric, Zakir
contents Large language models (LLMs) have exploded in popularity due to their ability to perform a wide array of natural language tasks. Text-based content moderation is one LLM use case that has received recent enthusiasm, however, there is little research investigating how LLMs perform in content moderation settings. In this work, we evaluate a suite of commodity LLMs on two common content moderation tasks: rule-based community moderation and toxic content detection. For rule-based community moderation, we instantiate 95 subcommunity specific LLMs by prompting GPT-3.5 with rules from 95 Reddit subcommunities. We find that GPT-3.5 is effective at rule-based moderation for many communities, achieving a median accuracy of 64% and a median precision of 83%. For toxicity detection, we evaluate a suite of commodity LLMs (GPT-3, GPT-3.5, GPT-4, Gemini Pro, LLAMA 2) and show that LLMs significantly outperform currently widespread toxicity classifiers. However, recent increases in model size add only marginal benefit to toxicity detection, suggesting a potential performance plateau for LLMs on toxicity detection tasks. We conclude by outlining avenues for future work in studying LLMs and content moderation.
format Preprint
id arxiv_https___arxiv_org_abs_2309_14517
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Watch Your Language: Investigating Content Moderation with Large Language Models
Kumar, Deepak
AbuHashem, Yousef
Durumeric, Zakir
Human-Computer Interaction
Artificial Intelligence
Computation and Language
Cryptography and Security
Social and Information Networks
Large language models (LLMs) have exploded in popularity due to their ability to perform a wide array of natural language tasks. Text-based content moderation is one LLM use case that has received recent enthusiasm, however, there is little research investigating how LLMs perform in content moderation settings. In this work, we evaluate a suite of commodity LLMs on two common content moderation tasks: rule-based community moderation and toxic content detection. For rule-based community moderation, we instantiate 95 subcommunity specific LLMs by prompting GPT-3.5 with rules from 95 Reddit subcommunities. We find that GPT-3.5 is effective at rule-based moderation for many communities, achieving a median accuracy of 64% and a median precision of 83%. For toxicity detection, we evaluate a suite of commodity LLMs (GPT-3, GPT-3.5, GPT-4, Gemini Pro, LLAMA 2) and show that LLMs significantly outperform currently widespread toxicity classifiers. However, recent increases in model size add only marginal benefit to toxicity detection, suggesting a potential performance plateau for LLMs on toxicity detection tasks. We conclude by outlining avenues for future work in studying LLMs and content moderation.
title Watch Your Language: Investigating Content Moderation with Large Language Models
topic Human-Computer Interaction
Artificial Intelligence
Computation and Language
Cryptography and Security
Social and Information Networks
url https://arxiv.org/abs/2309.14517