Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Daman Deep, Bhattacharjee, Ramanuj, Chakraborty, Abhijnan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918059874189312
author Singh, Daman Deep
Bhattacharjee, Ramanuj
Chakraborty, Abhijnan
author_facet Singh, Daman Deep
Bhattacharjee, Ramanuj
Chakraborty, Abhijnan
contents Hate speech detection across contemporary social media presents unique challenges due to linguistic diversity and the informal nature of online discourse. These challenges are further amplified in settings involving code-mixing, transliteration, and culturally nuanced expressions. While fine-tuned transformer models, such as BERT, have become standard for this task, we argue that recent large language models (LLMs) not only surpass them but also redefine the landscape of hate speech detection more broadly. To support this claim, we introduce IndoHateMix, a diverse, high-quality dataset capturing Hindi-English code-mixing and transliteration in the Indian context, providing a realistic benchmark to evaluate model robustness in complex multilingual scenarios where existing NLP methods often struggle. Our extensive experiments show that cutting-edge LLMs (such as LLaMA-3.1) consistently outperform task-specific BERT-based models, even when fine-tuned on significantly less data. With their superior generalization and adaptability, LLMs offer a transformative approach to mitigating online hate in diverse environments. This raises the question of whether future works should prioritize developing specialized models or focus on curating richer and more varied datasets to further enhance the effectiveness of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12744
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
Singh, Daman Deep
Bhattacharjee, Ramanuj
Chakraborty, Abhijnan
Computation and Language
Computers and Society
Hate speech detection across contemporary social media presents unique challenges due to linguistic diversity and the informal nature of online discourse. These challenges are further amplified in settings involving code-mixing, transliteration, and culturally nuanced expressions. While fine-tuned transformer models, such as BERT, have become standard for this task, we argue that recent large language models (LLMs) not only surpass them but also redefine the landscape of hate speech detection more broadly. To support this claim, we introduce IndoHateMix, a diverse, high-quality dataset capturing Hindi-English code-mixing and transliteration in the Indian context, providing a realistic benchmark to evaluate model robustness in complex multilingual scenarios where existing NLP methods often struggle. Our extensive experiments show that cutting-edge LLMs (such as LLaMA-3.1) consistently outperform task-specific BERT-based models, even when fine-tuned on significantly less data. With their superior generalization and adaptability, LLMs offer a transformative approach to mitigating online hate in diverse environments. This raises the question of whether future works should prioritize developing specialized models or focus on curating richer and more varied datasets to further enhance the effectiveness of LLMs.
title Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2506.12744