Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gajewska, Ewelina, Wawer, Michal, Budzynska, Katarzyna, Chudziak, Jaroslaw A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging a Multi-Agent LLM-Based System to Educate Teachers in Hate Incidents Management
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2025)
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2025)
Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2026)
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2026)
Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2025)
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2025)
Predicting the winner of the US 2024 elections using trust analytics
von: Budzynska, Katarzyna, et al.
Veröffentlicht: (2024)
von: Budzynska, Katarzyna, et al.
Veröffentlicht: (2024)
On Theoretically-Driven LLM Agents for Multi-Dimensional Discourse Analysis
von: Uberna, Maciej, et al.
Veröffentlicht: (2026)
von: Uberna, Maciej, et al.
Veröffentlicht: (2026)
ElliottAgents: A Natural Language-Driven Multi-Agent System for Stock Market Analysis and Prediction
von: Chudziak, Jarosław A., et al.
Veröffentlicht: (2025)
von: Chudziak, Jarosław A., et al.
Veröffentlicht: (2025)
Integrating Traditional Technical Analysis with AI: A Multi-Agent LLM-Based Approach to Stock Market Forecasting
von: Wawer, Michał, et al.
Veröffentlicht: (2025)
von: Wawer, Michał, et al.
Veröffentlicht: (2025)
The Lovelace Test of Intelligence: Can Humans Recognise and Esteem AI-Generated Art?
von: Gajewska, Ewelina
Veröffentlicht: (2025)
von: Gajewska, Ewelina
Veröffentlicht: (2025)
When AI Agents Disagree Like Humans: Reasoning Trace Analysis for Human-AI Collaborative Moderation
von: Wawer, Michał, et al.
Veröffentlicht: (2026)
von: Wawer, Michał, et al.
Veröffentlicht: (2026)
Ethos and Pathos in Online Group Discussions: Corpora for Polarisation Issues in Social Media
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2024)
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2024)
Heterogeneous Debate Engine: Identity-Grounded Cognitive Architecture for Resilient LLM-Based Ethical Tutoring
von: Masłowski, Jakub, et al.
Veröffentlicht: (2026)
von: Masłowski, Jakub, et al.
Veröffentlicht: (2026)
On Verifiable Legal Reasoning: A Multi-Agent Framework with Formalized Knowledge Representations
von: Sadowski, Albert, et al.
Veröffentlicht: (2025)
von: Sadowski, Albert, et al.
Veröffentlicht: (2025)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
TACLA: An LLM-Based Multi-Agent Tool for Transactional Analysis Training in Education
von: Zamojska, Monika, et al.
Veröffentlicht: (2025)
von: Zamojska, Monika, et al.
Veröffentlicht: (2025)
Multi-Agent Dialectical Refinement for Enhanced Argument Classification
von: Bąba, Jakub, et al.
Veröffentlicht: (2026)
von: Bąba, Jakub, et al.
Veröffentlicht: (2026)
GAIus: Combining Genai with Legal Clauses Retrieval for Knowledge-based Assistant
von: Matak, Michał, et al.
Veröffentlicht: (2025)
von: Matak, Michał, et al.
Veröffentlicht: (2025)
LLM-based Semantic Augmentation for Harmful Content Detection
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
A Natural Language Agentic Approach to Study Affective Polarization
von: Malvicini, Stephanie Anneris, et al.
Veröffentlicht: (2026)
von: Malvicini, Stephanie Anneris, et al.
Veröffentlicht: (2026)
Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation
von: Herrmann, Nils A., et al.
Veröffentlicht: (2026)
von: Herrmann, Nils A., et al.
Veröffentlicht: (2026)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Critical Challenges in Content Moderation for People Who Use Drugs (PWUD): Insights into Online Harm Reduction Practices from Moderators
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
Longitudinal Monitoring of LLM Content Moderation of Social Issues
von: Dai, Yunlang, et al.
Veröffentlicht: (2025)
von: Dai, Yunlang, et al.
Veröffentlicht: (2025)
From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
Harmful Suicide Content Detection
von: Park, Kyumin, et al.
Veröffentlicht: (2024)
von: Park, Kyumin, et al.
Veröffentlicht: (2024)
Identity-related Speech Suppression in Generative AI Content Moderation
von: Proebsting, Grace, et al.
Veröffentlicht: (2024)
von: Proebsting, Grace, et al.
Veröffentlicht: (2024)
Explainable Rule Application via Structured Prompting: A Neural-Symbolic Approach
von: Sadowski, Albert, et al.
Veröffentlicht: (2025)
von: Sadowski, Albert, et al.
Veröffentlicht: (2025)
Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models
von: Bonagiri, Akash, et al.
Veröffentlicht: (2025)
von: Bonagiri, Akash, et al.
Veröffentlicht: (2025)
Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
von: Maurer, Maximilian, et al.
Veröffentlicht: (2026)
von: Maurer, Maximilian, et al.
Veröffentlicht: (2026)
What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text
von: Kanepajs, Arturs, et al.
Veröffentlicht: (2025)
von: Kanepajs, Arturs, et al.
Veröffentlicht: (2025)
Conversational Agents to Facilitate Deliberation on Harmful Content in WhatsApp Groups
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2024)
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2024)
Who Decides in AI-Mediated Learning? The Agency Allocation Framework
von: Borchers, Conrad, et al.
Veröffentlicht: (2026)
von: Borchers, Conrad, et al.
Veröffentlicht: (2026)
Do Large Language Models Get Caught in Hofstadter-Mobius Loops?
von: Hryszko, Jaroslaw
Veröffentlicht: (2026)
von: Hryszko, Jaroslaw
Veröffentlicht: (2026)
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
von: Ghosh, Shaona, et al.
Veröffentlicht: (2024)
von: Ghosh, Shaona, et al.
Veröffentlicht: (2024)
LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models
von: Gupta, Sammriddh, et al.
Veröffentlicht: (2025)
von: Gupta, Sammriddh, et al.
Veröffentlicht: (2025)
Careless Whisper: Speech-to-Text Hallucination Harms
von: Koenecke, Allison, et al.
Veröffentlicht: (2024)
von: Koenecke, Allison, et al.
Veröffentlicht: (2024)
Taxonomizing Representational Harms using Speech Act Theory
von: Corvi, Emily, et al.
Veröffentlicht: (2025)
von: Corvi, Emily, et al.
Veröffentlicht: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
AI Content Moderation in Therapy Conversations
von: Kim, Jiwon, et al.
Veröffentlicht: (2026)
von: Kim, Jiwon, et al.
Veröffentlicht: (2026)
A Capabilities Approach to Studying Bias and Harm in Language Technologies
von: Nigatu, Hellina Hailu, et al.
Veröffentlicht: (2024)
von: Nigatu, Hellina Hailu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Leveraging a Multi-Agent LLM-Based System to Educate Teachers in Hate Incidents Management
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2025) -
Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2026) -
Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2025) -
Predicting the winner of the US 2024 elections using trust analytics
von: Budzynska, Katarzyna, et al.
Veröffentlicht: (2024) -
On Theoretically-Driven LLM Agents for Multi-Dimensional Discourse Analysis
von: Uberna, Maciej, et al.
Veröffentlicht: (2026)