HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rizwan, Naquee, Yimam, Seid Muhie, Dementieva, Daryna, Skupin, Florian, Fischer, Tim, Moskovskiy, Daniil, Borkar, Aarushi Ajay, Geislinger, Robert, Saha, Punyajoy, Roy, Sarthak, Semmann, Martin, Panchenko, Alexander, Biemann, Chris, Mukherjee, Animesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911041671135232
author Rizwan, Naquee
Yimam, Seid Muhie
Dementieva, Daryna
Skupin, Florian
Fischer, Tim
Moskovskiy, Daniil
Borkar, Aarushi Ajay
Geislinger, Robert
Saha, Punyajoy
Roy, Sarthak
Semmann, Martin
Panchenko, Alexander
Biemann, Chris
Mukherjee, Animesh
author_facet Rizwan, Naquee
Yimam, Seid Muhie
Dementieva, Daryna
Skupin, Florian
Fischer, Tim
Moskovskiy, Daniil
Borkar, Aarushi Ajay
Geislinger, Robert
Saha, Punyajoy
Roy, Sarthak
Semmann, Martin
Panchenko, Alexander
Biemann, Chris
Mukherjee, Animesh
contents Despite regulations imposed by nations and social media platforms, e.g. (Government of India, 2021; European Parliament and Council of the European Union, 2022), inter alia, hateful content persists as a significant challenge. Existing approaches primarily rely on reactive measures such as blocking or suspending offensive messages, with emerging strategies focusing on proactive measurements like detoxification and counterspeech. In our work, which we call HatePRISM, we conduct a comprehensive examination of hate speech regulations and strategies from three perspectives: country regulations, social platform policies, and NLP research datasets. Our findings reveal significant inconsistencies in hate speech definitions and moderation practices across jurisdictions and platforms, alongside a lack of alignment with research efforts. Based on these insights, we suggest ideas and research direction for further exploration of a unified framework for automated hate speech moderation incorporating diverse strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2507_04350
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation
Rizwan, Naquee
Yimam, Seid Muhie
Dementieva, Daryna
Skupin, Florian
Fischer, Tim
Moskovskiy, Daniil
Borkar, Aarushi Ajay
Geislinger, Robert
Saha, Punyajoy
Roy, Sarthak
Semmann, Martin
Panchenko, Alexander
Biemann, Chris
Mukherjee, Animesh
Computation and Language
Despite regulations imposed by nations and social media platforms, e.g. (Government of India, 2021; European Parliament and Council of the European Union, 2022), inter alia, hateful content persists as a significant challenge. Existing approaches primarily rely on reactive measures such as blocking or suspending offensive messages, with emerging strategies focusing on proactive measurements like detoxification and counterspeech. In our work, which we call HatePRISM, we conduct a comprehensive examination of hate speech regulations and strategies from three perspectives: country regulations, social platform policies, and NLP research datasets. Our findings reveal significant inconsistencies in hate speech definitions and moderation practices across jurisdictions and platforms, alongside a lack of alignment with research efforts. Based on these insights, we suggest ideas and research direction for further exploration of a unified framework for automated hate speech moderation incorporating diverse strategies.
title HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation
topic Computation and Language
url https://arxiv.org/abs/2507.04350