DefVerify: Do Hate Speech Models Reflect Their Dataset's Definition?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khurana, Urja, Nalisnick, Eric, Fokkens, Antske
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910781711319040
author Khurana, Urja
Nalisnick, Eric
Fokkens, Antske
author_facet Khurana, Urja
Nalisnick, Eric
Fokkens, Antske
contents When building a predictive model, it is often difficult to ensure that application-specific requirements are encoded by the model that will eventually be deployed. Consider researchers working on hate speech detection. They will have an idea of what is considered hate speech, but building a model that reflects their view accurately requires preserving those ideals throughout the workflow of data set construction and model training. Complications such as sampling bias, annotation bias, and model misspecification almost always arise, possibly resulting in a gap between the application specification and the model's actual behavior upon deployment. To address this issue for hate speech detection, we propose DefVerify: a 3-step procedure that (i) encodes a user-specified definition of hate speech, (ii) quantifies to what extent the model reflects the intended definition, and (iii) tries to identify the point of failure in the workflow. We use DefVerify to find gaps between definition and model behavior when applied to six popular hate speech benchmark datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15911
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DefVerify: Do Hate Speech Models Reflect Their Dataset's Definition?
Khurana, Urja
Nalisnick, Eric
Fokkens, Antske
Computation and Language
When building a predictive model, it is often difficult to ensure that application-specific requirements are encoded by the model that will eventually be deployed. Consider researchers working on hate speech detection. They will have an idea of what is considered hate speech, but building a model that reflects their view accurately requires preserving those ideals throughout the workflow of data set construction and model training. Complications such as sampling bias, annotation bias, and model misspecification almost always arise, possibly resulting in a gap between the application specification and the model's actual behavior upon deployment. To address this issue for hate speech detection, we propose DefVerify: a 3-step procedure that (i) encodes a user-specified definition of hate speech, (ii) quantifies to what extent the model reflects the intended definition, and (iii) tries to identify the point of failure in the workflow. We use DefVerify to find gaps between definition and model behavior when applied to six popular hate speech benchmark datasets.
title DefVerify: Do Hate Speech Models Reflect Their Dataset's Definition?
topic Computation and Language
url https://arxiv.org/abs/2410.15911