PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
Fuente:
arXiv
Saved in:
| Main Authors: | Kachwala, Zoher, Truong, Bao Tran, Muralidharan, Rasika, Kwak, Haewoon, An, Jisun, Menczer, Filippo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rematch: Robust and Efficient Matching of Local Knowledge Graphs to Improve Structural and Semantic Similarity
by: Kachwala, Zoher, et al.
Published: (2024)
by: Kachwala, Zoher, et al.
Published: (2024)
Can Lessons From Human Teams Be Applied to Multi-Agent Systems? The Role of Structure, Diversity, and Interaction Dynamics
by: Muralidharan, Rasika, et al.
Published: (2025)
by: Muralidharan, Rasika, et al.
Published: (2025)
Prefill-Guided Thinking for zero-shot detection of AI-generated images
by: Kachwala, Zoher, et al.
Published: (2025)
by: Kachwala, Zoher, et al.
Published: (2025)
LLMs Can Infer Political Alignment from Online Conversations
by: Lee, Byunghwee, et al.
Published: (2026)
by: Lee, Byunghwee, et al.
Published: (2026)
Somatic in the East, Psychological in the West?: Investigating Clinically-Grounded Cross-Cultural Depression Symptom Expression in LLMs
by: Sakai, Shintaro, et al.
Published: (2025)
by: Sakai, Shintaro, et al.
Published: (2025)
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
by: Huang, Fan, et al.
Published: (2026)
by: Huang, Fan, et al.
Published: (2026)
ToBlend: Token-Level Blending With an Ensemble of LLMs to Attack AI-Generated Text Detection
by: Huang, Fan, et al.
Published: (2024)
by: Huang, Fan, et al.
Published: (2024)
Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability
by: Huang, Fan, et al.
Published: (2026)
by: Huang, Fan, et al.
Published: (2026)
A semantic embedding space based on large language models for modelling human beliefs
by: Lee, Byunghwee, et al.
Published: (2024)
by: Lee, Byunghwee, et al.
Published: (2024)
ChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?
by: Huang, Fan, et al.
Published: (2024)
by: Huang, Fan, et al.
Published: (2024)
XChoice: Explainable Evaluation of AI-Human Alignment in LLM-based Constrained Choice Decision Making
by: Qi, Weihong, et al.
Published: (2026)
by: Qi, Weihong, et al.
Published: (2026)
Can we trust the evaluation on ChatGPT?
by: Aiyappa, Rachith, et al.
Published: (2023)
by: Aiyappa, Rachith, et al.
Published: (2023)
Quantifying Gender Stereotypes in Japan between 1900 and 1999 with Word Embeddings
by: Sakai, Shintaro, et al.
Published: (2025)
by: Sakai, Shintaro, et al.
Published: (2025)
Benchmarking zero-shot stance detection with FlanT5-XXL: Insights from training data, prompting, and decoding strategies into its near-SoTA performance
by: Aiyappa, Rachith, et al.
Published: (2024)
by: Aiyappa, Rachith, et al.
Published: (2024)
A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies
by: Qi, Weihong, et al.
Published: (2025)
by: Qi, Weihong, et al.
Published: (2025)
Community Moderation and the New Epistemology of Fact Checking on Social Media
by: Augenstein, Isabelle, et al.
Published: (2025)
by: Augenstein, Isabelle, et al.
Published: (2025)
Quantifying the Vulnerabilities of the Online Public Square to Adversarial Manipulation Tactics
by: Truong, Bao Tran, et al.
Published: (2019)
by: Truong, Bao Tran, et al.
Published: (2019)
Accuracy and Political Bias of News Source Credibility Ratings by Large Language Models
by: Yang, Kai-Cheng, et al.
Published: (2023)
by: Yang, Kai-Cheng, et al.
Published: (2023)
Community Notes are Vulnerable to Rater Bias and Manipulation
by: Truong, Bao Tran, et al.
Published: (2025)
by: Truong, Bao Tran, et al.
Published: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
by: Jiang, Han, et al.
Published: (2025)
by: Jiang, Han, et al.
Published: (2025)
Safeguarding Decentralized Social Media: LLM Agents for Automating Community Rule Compliance
by: La Cava, Lucio, et al.
Published: (2024)
by: La Cava, Lucio, et al.
Published: (2024)
Identifying Constructive Conflict in Online Discussions through Controversial yet Toxicity Resilient Posts
by: Seckin, Ozgur Can, et al.
Published: (2025)
by: Seckin, Ozgur Can, et al.
Published: (2025)
Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation
by: Samory, Mattia, et al.
Published: (2025)
by: Samory, Mattia, et al.
Published: (2025)
LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
by: Bojic, Ljubisa, et al.
Published: (2026)
by: Bojic, Ljubisa, et al.
Published: (2026)
Social Bias in Popular Question-Answering Benchmarks
by: Kraft, Angelie, et al.
Published: (2025)
by: Kraft, Angelie, et al.
Published: (2025)
Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models
by: Bonagiri, Akash, et al.
Published: (2025)
by: Bonagiri, Akash, et al.
Published: (2025)
Anatomy of an AI-powered malicious social botnet
by: Yang, Kai-Cheng, et al.
Published: (2023)
by: Yang, Kai-Cheng, et al.
Published: (2023)
What Helps Language Models Predict Human Beliefs: Demographics or Prior Stances?
by: Malone, Joseph, et al.
Published: (2025)
by: Malone, Joseph, et al.
Published: (2025)
Polarized Patterns of Language Toxicity and Sentiment of Debunking Posts on Social Media
by: Xu, Wentao, et al.
Published: (2025)
by: Xu, Wentao, et al.
Published: (2025)
A Marketplace for AI-Generated Adult Content and Deepfakes
by: Ghosh, Shalmoli, et al.
Published: (2026)
by: Ghosh, Shalmoli, et al.
Published: (2026)
Fine-Grained Behavior Simulation with Role-Playing Large Language Model on Social Media
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Understanding Mental Health Content on Social Media and Its Effect Towards Suicidal Ideation
by: Bhuiyan, Mohaiminul Islam, et al.
Published: (2025)
by: Bhuiyan, Mohaiminul Islam, et al.
Published: (2025)
Detection of Suicidal Risk on Social Media: A Hybrid Model
by: Yang, Zaihan, et al.
Published: (2025)
by: Yang, Zaihan, et al.
Published: (2025)
Rules, Cases, and Reasoning: Positivist Legal Theory as a Framework for Pluralistic AI Alignment
by: Caputo, Nicholas A.
Published: (2024)
by: Caputo, Nicholas A.
Published: (2024)
Why They Disagree: Decoding Differences in Opinions about AI Risk on the Lex Fridman Podcast
by: Truong, Nghi, et al.
Published: (2025)
by: Truong, Nghi, et al.
Published: (2025)
NoisyHate: Mining Online Human-Written Perturbations for Realistic Robustness Benchmarking of Content Moderation Models
by: Ye, Yiran, et al.
Published: (2023)
by: Ye, Yiran, et al.
Published: (2023)
Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search
by: DeVerna, Matthew R., et al.
Published: (2025)
by: DeVerna, Matthew R., et al.
Published: (2025)
Sentiment Analysis of Cyberbullying Data in Social Media
by: Susmitha, Arvapalli Sai, et al.
Published: (2024)
by: Susmitha, Arvapalli Sai, et al.
Published: (2024)
Enhanced Suicidal Ideation Detection from Social Media Using a CNN-BiLSTM Hybrid Model
by: Bhuiyan, Mohaiminul Islam, et al.
Published: (2025)
by: Bhuiyan, Mohaiminul Islam, et al.
Published: (2025)
Smart Trial: Evaluating the Use of Large Language Models for Recruiting Clinical Trial Participants via Social Media
by: Zhou, Xiaofan, et al.
Published: (2025)
by: Zhou, Xiaofan, et al.
Published: (2025)
Similar Items
-
Rematch: Robust and Efficient Matching of Local Knowledge Graphs to Improve Structural and Semantic Similarity
by: Kachwala, Zoher, et al.
Published: (2024) -
Can Lessons From Human Teams Be Applied to Multi-Agent Systems? The Role of Structure, Diversity, and Interaction Dynamics
by: Muralidharan, Rasika, et al.
Published: (2025) -
Prefill-Guided Thinking for zero-shot detection of AI-generated images
by: Kachwala, Zoher, et al.
Published: (2025) -
LLMs Can Infer Political Alignment from Online Conversations
by: Lee, Byunghwee, et al.
Published: (2026) -
Somatic in the East, Psychological in the West?: Investigating Clinically-Grounded Cross-Cultural Depression Symptom Expression in LLMs
by: Sakai, Shintaro, et al.
Published: (2025)