Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Gajewska, Ewelina, Wawer, Michal, Budzynska, Katarzyna, Chudziak, Jaroslaw A. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging a Multi-Agent LLM-Based System to Educate Teachers in Hate Incidents Management
by: Gajewska, Ewelina, et al.
Published: (2025)
by: Gajewska, Ewelina, et al.
Published: (2025)
Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
by: Gajewska, Ewelina, et al.
Published: (2026)
by: Gajewska, Ewelina, et al.
Published: (2026)
Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
by: Gajewska, Ewelina, et al.
Published: (2025)
by: Gajewska, Ewelina, et al.
Published: (2025)
Predicting the winner of the US 2024 elections using trust analytics
by: Budzynska, Katarzyna, et al.
Published: (2024)
by: Budzynska, Katarzyna, et al.
Published: (2024)
On Theoretically-Driven LLM Agents for Multi-Dimensional Discourse Analysis
by: Uberna, Maciej, et al.
Published: (2026)
by: Uberna, Maciej, et al.
Published: (2026)
ElliottAgents: A Natural Language-Driven Multi-Agent System for Stock Market Analysis and Prediction
by: Chudziak, Jarosław A., et al.
Published: (2025)
by: Chudziak, Jarosław A., et al.
Published: (2025)
Integrating Traditional Technical Analysis with AI: A Multi-Agent LLM-Based Approach to Stock Market Forecasting
by: Wawer, Michał, et al.
Published: (2025)
by: Wawer, Michał, et al.
Published: (2025)
The Lovelace Test of Intelligence: Can Humans Recognise and Esteem AI-Generated Art?
by: Gajewska, Ewelina
Published: (2025)
by: Gajewska, Ewelina
Published: (2025)
When AI Agents Disagree Like Humans: Reasoning Trace Analysis for Human-AI Collaborative Moderation
by: Wawer, Michał, et al.
Published: (2026)
by: Wawer, Michał, et al.
Published: (2026)
Ethos and Pathos in Online Group Discussions: Corpora for Polarisation Issues in Social Media
by: Gajewska, Ewelina, et al.
Published: (2024)
by: Gajewska, Ewelina, et al.
Published: (2024)
Heterogeneous Debate Engine: Identity-Grounded Cognitive Architecture for Resilient LLM-Based Ethical Tutoring
by: Masłowski, Jakub, et al.
Published: (2026)
by: Masłowski, Jakub, et al.
Published: (2026)
On Verifiable Legal Reasoning: A Multi-Agent Framework with Formalized Knowledge Representations
by: Sadowski, Albert, et al.
Published: (2025)
by: Sadowski, Albert, et al.
Published: (2025)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
TACLA: An LLM-Based Multi-Agent Tool for Transactional Analysis Training in Education
by: Zamojska, Monika, et al.
Published: (2025)
by: Zamojska, Monika, et al.
Published: (2025)
Multi-Agent Dialectical Refinement for Enhanced Argument Classification
by: Bąba, Jakub, et al.
Published: (2026)
by: Bąba, Jakub, et al.
Published: (2026)
GAIus: Combining Genai with Legal Clauses Retrieval for Knowledge-based Assistant
by: Matak, Michał, et al.
Published: (2025)
by: Matak, Michał, et al.
Published: (2025)
LLM-based Semantic Augmentation for Harmful Content Detection
by: Meguellati, Elyas, et al.
Published: (2025)
by: Meguellati, Elyas, et al.
Published: (2025)
HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
by: Wang, Kaixuan, et al.
Published: (2025)
by: Wang, Kaixuan, et al.
Published: (2025)
A Natural Language Agentic Approach to Study Affective Polarization
by: Malvicini, Stephanie Anneris, et al.
Published: (2026)
by: Malvicini, Stephanie Anneris, et al.
Published: (2026)
Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation
by: Herrmann, Nils A., et al.
Published: (2026)
by: Herrmann, Nils A., et al.
Published: (2026)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Critical Challenges in Content Moderation for People Who Use Drugs (PWUD): Insights into Online Harm Reduction Practices from Moderators
by: Wang, Kaixuan, et al.
Published: (2025)
by: Wang, Kaixuan, et al.
Published: (2025)
Longitudinal Monitoring of LLM Content Moderation of Social Issues
by: Dai, Yunlang, et al.
Published: (2025)
by: Dai, Yunlang, et al.
Published: (2025)
From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Harmful Suicide Content Detection
by: Park, Kyumin, et al.
Published: (2024)
by: Park, Kyumin, et al.
Published: (2024)
Identity-related Speech Suppression in Generative AI Content Moderation
by: Proebsting, Grace, et al.
Published: (2024)
by: Proebsting, Grace, et al.
Published: (2024)
Explainable Rule Application via Structured Prompting: A Neural-Symbolic Approach
by: Sadowski, Albert, et al.
Published: (2025)
by: Sadowski, Albert, et al.
Published: (2025)
Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models
by: Bonagiri, Akash, et al.
Published: (2025)
by: Bonagiri, Akash, et al.
Published: (2025)
Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
by: Maurer, Maximilian, et al.
Published: (2026)
by: Maurer, Maximilian, et al.
Published: (2026)
What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text
by: Kanepajs, Arturs, et al.
Published: (2025)
by: Kanepajs, Arturs, et al.
Published: (2025)
Conversational Agents to Facilitate Deliberation on Harmful Content in WhatsApp Groups
by: Agarwal, Dhruv, et al.
Published: (2024)
by: Agarwal, Dhruv, et al.
Published: (2024)
Who Decides in AI-Mediated Learning? The Agency Allocation Framework
by: Borchers, Conrad, et al.
Published: (2026)
by: Borchers, Conrad, et al.
Published: (2026)
Do Large Language Models Get Caught in Hofstadter-Mobius Loops?
by: Hryszko, Jaroslaw
Published: (2026)
by: Hryszko, Jaroslaw
Published: (2026)
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
by: Ghosh, Shaona, et al.
Published: (2024)
by: Ghosh, Shaona, et al.
Published: (2024)
LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models
by: Gupta, Sammriddh, et al.
Published: (2025)
by: Gupta, Sammriddh, et al.
Published: (2025)
Careless Whisper: Speech-to-Text Hallucination Harms
by: Koenecke, Allison, et al.
Published: (2024)
by: Koenecke, Allison, et al.
Published: (2024)
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025)
by: Corvi, Emily, et al.
Published: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
AI Content Moderation in Therapy Conversations
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
A Capabilities Approach to Studying Bias and Harm in Language Technologies
by: Nigatu, Hellina Hailu, et al.
Published: (2024)
by: Nigatu, Hellina Hailu, et al.
Published: (2024)
Similar Items
-
Leveraging a Multi-Agent LLM-Based System to Educate Teachers in Hate Incidents Management
by: Gajewska, Ewelina, et al.
Published: (2025) -
Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
by: Gajewska, Ewelina, et al.
Published: (2026) -
Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
by: Gajewska, Ewelina, et al.
Published: (2025) -
Predicting the winner of the US 2024 elections using trust analytics
by: Budzynska, Katarzyna, et al.
Published: (2024) -
On Theoretically-Driven LLM Agents for Multi-Dimensional Discourse Analysis
by: Uberna, Maciej, et al.
Published: (2026)