From Moderation to Mediation: Can LLMs Serve as Mediators in Online Flame Wars?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Dawei, Alnaibari, Abdullah, Bisharat, Arslan, Sandoval, Manny, Hall, Deborah, Silva, Yasin, Liu, Huan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915820847759360
author Li, Dawei
Alnaibari, Abdullah
Bisharat, Arslan
Sandoval, Manny
Hall, Deborah
Silva, Yasin
Liu, Huan
author_facet Li, Dawei
Alnaibari, Abdullah
Bisharat, Arslan
Sandoval, Manny
Hall, Deborah
Silva, Yasin
Liu, Huan
contents The rapid advancement of large language models (LLMs) has opened new possibilities for AI for good applications. As LLMs increasingly mediate online communication, their potential to foster empathy and constructive dialogue becomes an important frontier for responsible AI research. This work explores whether LLMs can serve not only as moderators that detect harmful content, but as mediators capable of understanding and de-escalating online conflicts. Our framework decomposes mediation into two subtasks: judgment, where an LLM evaluates the fairness and emotional dynamics of a conversation, and steering, where it generates empathetic, de-escalatory messages to guide participants toward resolution. To assess mediation quality, we construct a large Reddit-based dataset and propose a multi-stage evaluation pipeline combining principle-based scoring, user simulation, and human comparison. Experiments show that API-based models outperform open-source counterparts in both reasoning and intervention alignment when doing mediation. Our findings highlight both the promise and limitations of current LLMs as emerging agents for online social mediation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03005
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Moderation to Mediation: Can LLMs Serve as Mediators in Online Flame Wars?
Li, Dawei
Alnaibari, Abdullah
Bisharat, Arslan
Sandoval, Manny
Hall, Deborah
Silva, Yasin
Liu, Huan
Artificial Intelligence
The rapid advancement of large language models (LLMs) has opened new possibilities for AI for good applications. As LLMs increasingly mediate online communication, their potential to foster empathy and constructive dialogue becomes an important frontier for responsible AI research. This work explores whether LLMs can serve not only as moderators that detect harmful content, but as mediators capable of understanding and de-escalating online conflicts. Our framework decomposes mediation into two subtasks: judgment, where an LLM evaluates the fairness and emotional dynamics of a conversation, and steering, where it generates empathetic, de-escalatory messages to guide participants toward resolution. To assess mediation quality, we construct a large Reddit-based dataset and propose a multi-stage evaluation pipeline combining principle-based scoring, user simulation, and human comparison. Experiments show that API-based models outperform open-source counterparts in both reasoning and intervention alignment when doing mediation. Our findings highlight both the promise and limitations of current LLMs as emerging agents for online social mediation.
title From Moderation to Mediation: Can LLMs Serve as Mediators in Online Flame Wars?
topic Artificial Intelligence
url https://arxiv.org/abs/2512.03005