Evolving Hate Speech Online: An Adaptive Framework for Detection and Mitigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ali, Shiza, Blackburn, Jeremy, Stringhini, Gianluca
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909503906119680
author Ali, Shiza
Blackburn, Jeremy
Stringhini, Gianluca
author_facet Ali, Shiza
Blackburn, Jeremy
Stringhini, Gianluca
contents The proliferation of social media platforms has led to an increase in the spread of hate speech, particularly targeting vulnerable communities. Unfortunately, existing methods for automatically identifying and blocking toxic language rely on pre-constructed lexicons, making them reactive rather than adaptive. As such, these approaches become less effective over time, especially when new communities are targeted with slurs not included in the original datasets. To address this issue, we present an adaptive approach that uses word embeddings to update lexicons and develop a hybrid model that adjusts to emerging slurs and new linguistic patterns. This approach can effectively detect toxic language, including intentional spelling mistakes employed by aggressors to avoid detection. Our hybrid model, which combines BERT with lexicon-based techniques, achieves an accuracy of 95% for most state-of-the-art datasets. Our work has significant implications for creating safer online environments by improving the detection of toxic content and proactively updating the lexicon. Content Warning: This paper contains examples of hate speech that may be triggering.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evolving Hate Speech Online: An Adaptive Framework for Detection and Mitigation
Ali, Shiza
Blackburn, Jeremy
Stringhini, Gianluca
Computation and Language
Social and Information Networks
The proliferation of social media platforms has led to an increase in the spread of hate speech, particularly targeting vulnerable communities. Unfortunately, existing methods for automatically identifying and blocking toxic language rely on pre-constructed lexicons, making them reactive rather than adaptive. As such, these approaches become less effective over time, especially when new communities are targeted with slurs not included in the original datasets. To address this issue, we present an adaptive approach that uses word embeddings to update lexicons and develop a hybrid model that adjusts to emerging slurs and new linguistic patterns. This approach can effectively detect toxic language, including intentional spelling mistakes employed by aggressors to avoid detection. Our hybrid model, which combines BERT with lexicon-based techniques, achieves an accuracy of 95% for most state-of-the-art datasets. Our work has significant implications for creating safer online environments by improving the detection of toxic content and proactively updating the lexicon. Content Warning: This paper contains examples of hate speech that may be triggering.
title Evolving Hate Speech Online: An Adaptive Framework for Detection and Mitigation
topic Computation and Language
Social and Information Networks
url https://arxiv.org/abs/2502.10921