PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kachwala, Zoher, Truong, Bao Tran, Muralidharan, Rasika, Kwak, Haewoon, An, Jisun, Menczer, Filippo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911691720097792
author Kachwala, Zoher
Truong, Bao Tran
Muralidharan, Rasika
Kwak, Haewoon
An, Jisun
Menczer, Filippo
author_facet Kachwala, Zoher
Truong, Bao Tran
Muralidharan, Rasika
Kwak, Haewoon
An, Jisun
Menczer, Filippo
contents Social media are shifting towards pluralism -- community-governed platforms where groups define their own norms. What violates rules in one community may be perfectly acceptable in another. Can AI models help moderate such pluralistic communities? We formalize the task as a multiple-choice problem, mirroring how human moderators operate in the real world: given a comment and its surrounding context, identify which specific rule, if any, is violated. We introduce PluRule, a multimodal, multilingual benchmark for detecting 13,371 rule violations across 1,989 Reddit communities spanning 2,885 rules in 9 languages. Using this benchmark, we show that state-of-the-art vision-language models struggle significantly: even GPT-5.2 with high reasoning performs only slightly better than a trivial baseline. We also find that bigger models and increased context provide marginal gains, and universal rules like civility and self-promotion are easier to detect. Our results show that moderation of pluralistic communities on social media is a fundamental challenge for language models. Our code and benchmark are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17187
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
Kachwala, Zoher
Truong, Bao Tran
Muralidharan, Rasika
Kwak, Haewoon
An, Jisun
Menczer, Filippo
Computation and Language
Artificial Intelligence
Computers and Society
Social media are shifting towards pluralism -- community-governed platforms where groups define their own norms. What violates rules in one community may be perfectly acceptable in another. Can AI models help moderate such pluralistic communities? We formalize the task as a multiple-choice problem, mirroring how human moderators operate in the real world: given a comment and its surrounding context, identify which specific rule, if any, is violated. We introduce PluRule, a multimodal, multilingual benchmark for detecting 13,371 rule violations across 1,989 Reddit communities spanning 2,885 rules in 9 languages. Using this benchmark, we show that state-of-the-art vision-language models struggle significantly: even GPT-5.2 with high reasoning performs only slightly better than a trivial baseline. We also find that bigger models and increased context provide marginal gains, and universal rules like civility and self-promotion are easier to detect. Our results show that moderation of pluralistic communities on social media is a fundamental challenge for language models. Our code and benchmark are publicly available.
title PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2605.17187