Classification is a RAG problem: A case study on hate speech detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Willats, Richard, Pennington, Josh, Mohan, Aravind, Vidgen, Bertie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916887784325120
author Willats, Richard
Pennington, Josh
Mohan, Aravind
Vidgen, Bertie
author_facet Willats, Richard
Pennington, Josh
Mohan, Aravind
Vidgen, Bertie
contents Robust content moderation requires classification systems that can quickly adapt to evolving policies without costly retraining. We present classification using Retrieval-Augmented Generation (RAG), which shifts traditional classification tasks from determining the correct category in accordance with pre-trained parameters to evaluating content in relation to contextual knowledge retrieved at inference. In hate speech detection, this transforms the task from "is this hate speech?" to "does this violate the hate speech policy?" Our Contextual Policy Engine (CPE) - an agentic RAG system - demonstrates this approach and offers three key advantages: (1) robust classification accuracy comparable to leading commercial systems, (2) inherent explainability via retrieved policy segments, and (3) dynamic policy updates without model retraining. Through three experiments, we demonstrate strong baseline performance and show that the system can apply fine-grained policy control by correctly adjusting protection for specific identity groups without requiring retraining or compromising overall performance. These findings establish that RAG can transform classification into a more flexible, transparent, and adaptable process for content moderation and wider classification problems.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06204
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Classification is a RAG problem: A case study on hate speech detection
Willats, Richard
Pennington, Josh
Mohan, Aravind
Vidgen, Bertie
Computation and Language
Artificial Intelligence
Machine Learning
Robust content moderation requires classification systems that can quickly adapt to evolving policies without costly retraining. We present classification using Retrieval-Augmented Generation (RAG), which shifts traditional classification tasks from determining the correct category in accordance with pre-trained parameters to evaluating content in relation to contextual knowledge retrieved at inference. In hate speech detection, this transforms the task from "is this hate speech?" to "does this violate the hate speech policy?" Our Contextual Policy Engine (CPE) - an agentic RAG system - demonstrates this approach and offers three key advantages: (1) robust classification accuracy comparable to leading commercial systems, (2) inherent explainability via retrieved policy segments, and (3) dynamic policy updates without model retraining. Through three experiments, we demonstrate strong baseline performance and show that the system can apply fine-grained policy control by correctly adjusting protection for specific identity groups without requiring retraining or compromising overall performance. These findings establish that RAG can transform classification into a more flexible, transparent, and adaptable process for content moderation and wider classification problems.
title Classification is a RAG problem: A case study on hate speech detection
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.06204