HateBuffer: Safeguarding Content Moderators' Mental Well-Being through Hate Speech Content Modification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Subin, Kim, Jeonghyun, Choi, Jeanne, Seering, Joseph, Lee, Uichin, Lee, Sung-Ju
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916875130109952
author Park, Subin
Kim, Jeonghyun
Choi, Jeanne
Seering, Joseph
Lee, Uichin
Lee, Sung-Ju
author_facet Park, Subin
Kim, Jeonghyun
Choi, Jeanne
Seering, Joseph
Lee, Uichin
Lee, Sung-Ju
contents Hate speech remains a persistent and unresolved challenge in online platforms. Content moderators, working on the front lines to review user-generated content and shield viewers from hate speech, often find themselves unprotected from the mental burden as they continuously engage with offensive language. To safeguard moderators' mental well-being, we designed HateBuffer, which anonymizes targets of hate speech, paraphrases offensive expressions into less offensive forms, and shows the original expressions when moderators opt to see them. Our user study with 80 participants consisted of a simulated hate speech moderation task set on a fictional news platform, followed by semi-structured interviews. Although participants rated the hate severity of comments lower while using HateBuffer, contrary to our expectations, they did not experience improved emotion or reduced fatigue compared with the control group. In interviews, however, participants described HateBuffer as an effective buffer against emotional contagion and the normalization of biased opinions in hate speech. Notably, HateBuffer did not compromise moderation accuracy and even contributed to a slight increase in recall. We explore possible explanations for the discrepancy between the perceived benefits of HateBuffer and its measured impact on mental well-being. We also underscore the promise of text-based content modification techniques as tools for a healthier content moderation environment.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HateBuffer: Safeguarding Content Moderators' Mental Well-Being through Hate Speech Content Modification
Park, Subin
Kim, Jeonghyun
Choi, Jeanne
Seering, Joseph
Lee, Uichin
Lee, Sung-Ju
Human-Computer Interaction
Hate speech remains a persistent and unresolved challenge in online platforms. Content moderators, working on the front lines to review user-generated content and shield viewers from hate speech, often find themselves unprotected from the mental burden as they continuously engage with offensive language. To safeguard moderators' mental well-being, we designed HateBuffer, which anonymizes targets of hate speech, paraphrases offensive expressions into less offensive forms, and shows the original expressions when moderators opt to see them. Our user study with 80 participants consisted of a simulated hate speech moderation task set on a fictional news platform, followed by semi-structured interviews. Although participants rated the hate severity of comments lower while using HateBuffer, contrary to our expectations, they did not experience improved emotion or reduced fatigue compared with the control group. In interviews, however, participants described HateBuffer as an effective buffer against emotional contagion and the normalization of biased opinions in hate speech. Notably, HateBuffer did not compromise moderation accuracy and even contributed to a slight increase in recall. We explore possible explanations for the discrepancy between the perceived benefits of HateBuffer and its measured impact on mental well-being. We also underscore the promise of text-based content modification techniques as tools for a healthier content moderation environment.
title HateBuffer: Safeguarding Content Moderators' Mental Well-Being through Hate Speech Content Modification
topic Human-Computer Interaction
url https://arxiv.org/abs/2508.00439