Predictively Combatting Toxicity in Health-related Online Discussions through Machine Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paz-Ruza, Jorge, Alonso-Betanzos, Amparo, Guijarro-Berdiñas, Bertha, Eiras-Franco, Carlos
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908376040996864
author Paz-Ruza, Jorge
Alonso-Betanzos, Amparo
Guijarro-Berdiñas, Bertha
Eiras-Franco, Carlos
author_facet Paz-Ruza, Jorge
Alonso-Betanzos, Amparo
Guijarro-Berdiñas, Bertha
Eiras-Franco, Carlos
contents In health-related topics, user toxicity in online discussions frequently becomes a source of social conflict or promotion of dangerous, unscientific behaviour; common approaches for battling it include different forms of detection, flagging and/or removal of existing toxic comments, which is often counterproductive for platforms and users alike. In this work, we propose the alternative of combatting user toxicity predictively, anticipating where a user could interact toxically in health-related online discussions. Applying a Collaborative Filtering-based Machine Learning methodology, we predict the toxicity in COVID-related conversations between any user and subcommunity of Reddit, surpassing 80% predictive performance in relevant metrics, and allowing us to prevent the pairing of conflicting users and subcommunities.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17068
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Predictively Combatting Toxicity in Health-related Online Discussions through Machine Learning
Paz-Ruza, Jorge
Alonso-Betanzos, Amparo
Guijarro-Berdiñas, Bertha
Eiras-Franco, Carlos
Computation and Language
Machine Learning
Social and Information Networks
In health-related topics, user toxicity in online discussions frequently becomes a source of social conflict or promotion of dangerous, unscientific behaviour; common approaches for battling it include different forms of detection, flagging and/or removal of existing toxic comments, which is often counterproductive for platforms and users alike. In this work, we propose the alternative of combatting user toxicity predictively, anticipating where a user could interact toxically in health-related online discussions. Applying a Collaborative Filtering-based Machine Learning methodology, we predict the toxicity in COVID-related conversations between any user and subcommunity of Reddit, surpassing 80% predictive performance in relevant metrics, and allowing us to prevent the pairing of conflicting users and subcommunities.
title Predictively Combatting Toxicity in Health-related Online Discussions through Machine Learning
topic Computation and Language
Machine Learning
Social and Information Networks
url https://arxiv.org/abs/2505.17068