Improving Hate Speech Classification with Cross-Taxonomy Dataset Integration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fillies, Jan, Paschke, Adrian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910863549530112
author Fillies, Jan
Paschke, Adrian
author_facet Fillies, Jan
Paschke, Adrian
contents Algorithmic hate speech detection faces significant challenges due to the diverse definitions and datasets used in research and practice. Social media platforms, legal frameworks, and institutions each apply distinct yet overlapping definitions, complicating classification efforts. This study addresses these challenges by demonstrating that existing datasets and taxonomies can be integrated into a unified model, enhancing prediction performance and reducing reliance on multiple specialized classifiers. The work introduces a universal taxonomy and a hate speech classifier capable of detecting a wide range of definitions within a single framework. Our approach is validated by combining two widely used but differently annotated datasets, showing improved classification performance on an independent test set. This work highlights the potential of dataset and taxonomy integration in advancing hate speech detection, increasing efficiency, and ensuring broader applicability across contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2503_05357
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Hate Speech Classification with Cross-Taxonomy Dataset Integration
Fillies, Jan
Paschke, Adrian
Computation and Language
Artificial Intelligence
Machine Learning
Social and Information Networks
Algorithmic hate speech detection faces significant challenges due to the diverse definitions and datasets used in research and practice. Social media platforms, legal frameworks, and institutions each apply distinct yet overlapping definitions, complicating classification efforts. This study addresses these challenges by demonstrating that existing datasets and taxonomies can be integrated into a unified model, enhancing prediction performance and reducing reliance on multiple specialized classifiers. The work introduces a universal taxonomy and a hate speech classifier capable of detecting a wide range of definitions within a single framework. Our approach is validated by combining two widely used but differently annotated datasets, showing improved classification performance on an independent test set. This work highlights the potential of dataset and taxonomy integration in advancing hate speech detection, increasing efficiency, and ensuring broader applicability across contexts.
title Improving Hate Speech Classification with Cross-Taxonomy Dataset Integration
topic Computation and Language
Artificial Intelligence
Machine Learning
Social and Information Networks
url https://arxiv.org/abs/2503.05357