DefVerify: Do Hate Speech Models Reflect Their Dataset's Definition?
Fuente:
arXiv
Saved in:
| Main Authors: | Khurana, Urja, Nalisnick, Eric, Fokkens, Antske |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?
by: Khurana, Urja, et al.
Published: (2024)
by: Khurana, Urja, et al.
Published: (2024)
Not All Subjectivity Is the Same! Defining Desiderata for the Evaluation of Subjectivity in NLP
by: Khurana, Urja, et al.
Published: (2026)
by: Khurana, Urja, et al.
Published: (2026)
On the Low-Rank Parametrization of Reward Models for Controlled Language Generation
by: Troshin, Sergey, et al.
Published: (2024)
by: Troshin, Sergey, et al.
Published: (2024)
Investigating the Robustness of Modelling Decisions for Few-Shot Cross-Topic Stance Detection: A Preregistered Study
by: Reuver, Myrthe, et al.
Published: (2024)
by: Reuver, Myrthe, et al.
Published: (2024)
Balancing the Scales: Reinforcement Learning for Fair Classification
by: Eshuijs, Leon, et al.
Published: (2024)
by: Eshuijs, Leon, et al.
Published: (2024)
The Role of Syntactic Span Preferences in Post-Hoc Explanation Disagreement
by: Kamp, Jonathan, et al.
Published: (2024)
by: Kamp, Jonathan, et al.
Published: (2024)
Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
by: Kamp, Jonathan, et al.
Published: (2025)
by: Kamp, Jonathan, et al.
Published: (2025)
Improving Causal Interventions in Amnesic Probing with Mean Projection or LEACE
by: Dobrzeniecka, Alicja, et al.
Published: (2025)
by: Dobrzeniecka, Alicja, et al.
Published: (2025)
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
by: Eshuijs, Leon, et al.
Published: (2025)
by: Eshuijs, Leon, et al.
Published: (2025)
Asking a Language Model for Diverse Responses
by: Troshin, Sergey, et al.
Published: (2025)
by: Troshin, Sergey, et al.
Published: (2025)
DefAn: Definitive Answer Dataset for LLMs Hallucination Evaluation
by: Rahman, A B M Ashikur, et al.
Published: (2024)
by: Rahman, A B M Ashikur, et al.
Published: (2024)
SciDef: Automating Definition Extraction from Academic Literature with Large Language Models
by: Kučera, Filip, et al.
Published: (2026)
by: Kučera, Filip, et al.
Published: (2026)
Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
by: Troshin, Sergey, et al.
Published: (2025)
by: Troshin, Sergey, et al.
Published: (2025)
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
Untangling Hate Speech Definitions: A Semantic Componential Analysis Across Cultures and Domains
by: Korre, Katerina, et al.
Published: (2024)
by: Korre, Katerina, et al.
Published: (2024)
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2024)
by: Piot, Paloma, et al.
Published: (2024)
Active Learning for Multilingual Fingerspelling Corpora
by: Wang, Shuai, et al.
Published: (2023)
by: Wang, Shuai, et al.
Published: (2023)
Empirical Evaluation of Public HateSpeech Datasets
by: Jaf, Sadar, et al.
Published: (2024)
by: Jaf, Sadar, et al.
Published: (2024)
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages
by: Muhammad, Shamsuddeen Hassan, et al.
Published: (2025)
by: Muhammad, Shamsuddeen Hassan, et al.
Published: (2025)
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
by: Piot, Paloma, et al.
Published: (2024)
by: Piot, Paloma, et al.
Published: (2024)
ProvocationProbe: Instigating Hate Speech Dataset from Twitter
by: Kumar, Abhay, et al.
Published: (2024)
by: Kumar, Abhay, et al.
Published: (2024)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter
by: Tonneau, Manuel, et al.
Published: (2024)
by: Tonneau, Manuel, et al.
Published: (2024)
A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance
by: Melis, Matteo, et al.
Published: (2025)
by: Melis, Matteo, et al.
Published: (2025)
BIDWESH: A Bangla Regional Based Hate Speech Detection Dataset
by: Fayaz, Azizul Hakim, et al.
Published: (2025)
by: Fayaz, Azizul Hakim, et al.
Published: (2025)
HateDebias: On the Diversity and Variability of Hate Speech Debiasing
by: Wu, Hongyan, et al.
Published: (2024)
by: Wu, Hongyan, et al.
Published: (2024)
Towards Weakly-Supervised Hate Speech Classification Across Datasets
by: Jin, Yiping, et al.
Published: (2023)
by: Jin, Yiping, et al.
Published: (2023)
From Languages to Geographies: Towards Evaluating Cultural Bias in Hate Speech Datasets
by: Tonneau, Manuel, et al.
Published: (2024)
by: Tonneau, Manuel, et al.
Published: (2024)
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models
by: Bui, Minh Duc, et al.
Published: (2024)
by: Bui, Minh Duc, et al.
Published: (2024)
"Is Hate Lost in Translation?": Evaluation of Multilingual LGBTQIA+ Hate Speech Detection
by: Chan, Fai Leui, et al.
Published: (2024)
by: Chan, Fai Leui, et al.
Published: (2024)
Web(er) of Hate: A Survey on How Hate Speech Is Typed
by: Wang, Luna, et al.
Published: (2025)
by: Wang, Luna, et al.
Published: (2025)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
by: Almohaimeed, Saad, et al.
Published: (2025)
by: Almohaimeed, Saad, et al.
Published: (2025)
BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla
by: Haider, Fabiha, et al.
Published: (2024)
by: Haider, Fabiha, et al.
Published: (2024)
Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets
by: Cheremetiev, Vassiliy, et al.
Published: (2025)
by: Cheremetiev, Vassiliy, et al.
Published: (2025)
Diagnosing Hate Speech Classification: Where Do Humans and Machines Disagree, and Why?
by: Yang, Xilin
Published: (2024)
by: Yang, Xilin
Published: (2024)
BengaliSent140: A Large-Scale Bengali Binary Sentiment Dataset for Hate and Non-Hate Speech Classification
by: Islam, Akif, et al.
Published: (2026)
by: Islam, Akif, et al.
Published: (2026)
HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
by: Sen, Tanmay, et al.
Published: (2024)
by: Sen, Tanmay, et al.
Published: (2024)
The Unseen Targets of Hate -- A Systematic Review of Hateful Communication Datasets
by: Yu, Zehui, et al.
Published: (2024)
by: Yu, Zehui, et al.
Published: (2024)
Outcome-Constrained Large Language Models for Countering Hate Speech
by: Hong, Lingzi, et al.
Published: (2024)
by: Hong, Lingzi, et al.
Published: (2024)
Bangla Hate Speech Classification with Fine-tuned Transformer Models
by: Jafari, Yalda Keivan, et al.
Published: (2025)
by: Jafari, Yalda Keivan, et al.
Published: (2025)
Similar Items
-
Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?
by: Khurana, Urja, et al.
Published: (2024) -
Not All Subjectivity Is the Same! Defining Desiderata for the Evaluation of Subjectivity in NLP
by: Khurana, Urja, et al.
Published: (2026) -
On the Low-Rank Parametrization of Reward Models for Controlled Language Generation
by: Troshin, Sergey, et al.
Published: (2024) -
Investigating the Robustness of Modelling Decisions for Few-Shot Cross-Topic Stance Detection: A Preregistered Study
by: Reuver, Myrthe, et al.
Published: (2024) -
Balancing the Scales: Reinforcement Learning for Fair Classification
by: Eshuijs, Leon, et al.
Published: (2024)