Predictively Combatting Toxicity in Health-related Online Discussions through Machine Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908376040996864 |
|---|---|
| author | Paz-Ruza, Jorge Alonso-Betanzos, Amparo Guijarro-Berdiñas, Bertha Eiras-Franco, Carlos |
| author_facet | Paz-Ruza, Jorge Alonso-Betanzos, Amparo Guijarro-Berdiñas, Bertha Eiras-Franco, Carlos |
| contents | In health-related topics, user toxicity in online discussions frequently becomes a source of social conflict or promotion of dangerous, unscientific behaviour; common approaches for battling it include different forms of detection, flagging and/or removal of existing toxic comments, which is often counterproductive for platforms and users alike. In this work, we propose the alternative of combatting user toxicity predictively, anticipating where a user could interact toxically in health-related online discussions. Applying a Collaborative Filtering-based Machine Learning methodology, we predict the toxicity in COVID-related conversations between any user and subcommunity of Reddit, surpassing 80% predictive performance in relevant metrics, and allowing us to prevent the pairing of conflicting users and subcommunities. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_17068 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Predictively Combatting Toxicity in Health-related Online Discussions through Machine Learning Paz-Ruza, Jorge Alonso-Betanzos, Amparo Guijarro-Berdiñas, Bertha Eiras-Franco, Carlos Computation and Language Machine Learning Social and Information Networks In health-related topics, user toxicity in online discussions frequently becomes a source of social conflict or promotion of dangerous, unscientific behaviour; common approaches for battling it include different forms of detection, flagging and/or removal of existing toxic comments, which is often counterproductive for platforms and users alike. In this work, we propose the alternative of combatting user toxicity predictively, anticipating where a user could interact toxically in health-related online discussions. Applying a Collaborative Filtering-based Machine Learning methodology, we predict the toxicity in COVID-related conversations between any user and subcommunity of Reddit, surpassing 80% predictive performance in relevant metrics, and allowing us to prevent the pairing of conflicting users and subcommunities. |
| title | Predictively Combatting Toxicity in Health-related Online Discussions through Machine Learning |
| topic | Computation and Language Machine Learning Social and Information Networks |
| url | https://arxiv.org/abs/2505.17068 |