Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Piot, Paloma, Otero, David, Martín-Rodilla, Patricia, Parapar, Javier |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2024)
by: Piot, Paloma, et al.
Published: (2024)
Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
by: Piot, Paloma, et al.
Published: (2024)
by: Piot, Paloma, et al.
Published: (2024)
Towards Efficient and Explainable Hate Speech Detection via Model Distillation
by: Piot, Paloma, et al.
Published: (2024)
by: Piot, Paloma, et al.
Published: (2024)
WATCHED: A Web AI Agent Tool for Combating Hate Speech by Expanding Data
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Bridging Gaps in Hate Speech Detection: Meta-Collections and Benchmarks for Low-Resource Iberian Languages
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
by: Kidder, William, et al.
Published: (2024)
by: Kidder, William, et al.
Published: (2024)
xList-Hate: A Checklist-Based Framework for Interpretable and Generalizable Hate Speech Detection
by: Girón, Adrián, et al.
Published: (2026)
by: Girón, Adrián, et al.
Published: (2026)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
by: Giorgi, Tommaso, et al.
Published: (2024)
by: Giorgi, Tommaso, et al.
Published: (2024)
Fact or Fiction? Can LLMs be Reliable Annotators for Political Truths?
by: Chatrath, Veronica, et al.
Published: (2024)
by: Chatrath, Veronica, et al.
Published: (2024)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
by: Das, Amit, et al.
Published: (2024)
by: Das, Amit, et al.
Published: (2024)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
by: Proskurina, Irina, et al.
Published: (2025)
by: Proskurina, Irina, et al.
Published: (2025)
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
by: Ni, Jingwei, et al.
Published: (2024)
by: Ni, Jingwei, et al.
Published: (2024)
Empirical Evaluation of Public HateSpeech Datasets
by: Jaf, Sadar, et al.
Published: (2024)
by: Jaf, Sadar, et al.
Published: (2024)
TalkDep: Clinically Grounded LLM Personas for Conversation-Centric Depression Screening
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
SelectLLM: Can LLMs Select Important Instructions to Annotate?
by: Parkar, Ritik Sachin, et al.
Published: (2024)
by: Parkar, Ritik Sachin, et al.
Published: (2024)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
by: Hu, Yujia, et al.
Published: (2026)
by: Hu, Yujia, et al.
Published: (2026)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
by: Lee, Nayeon, et al.
Published: (2023)
by: Lee, Nayeon, et al.
Published: (2023)
Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection
by: Kim, Yumin, et al.
Published: (2025)
by: Kim, Yumin, et al.
Published: (2025)
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025)
by: Bauer, Julie, et al.
Published: (2025)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
by: Kumarage, Tharindu, et al.
Published: (2024)
by: Kumarage, Tharindu, et al.
Published: (2024)
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
by: Bajpai, Prasoon, et al.
Published: (2024)
by: Bajpai, Prasoon, et al.
Published: (2024)
Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
by: Zhou, Tianyi, et al.
Published: (2025)
by: Zhou, Tianyi, et al.
Published: (2025)
Leveraging Weakly Annotated Data for Hate Speech Detection in Code-Mixed Hinglish: A Feasibility-Driven Transfer Learning Approach with Large Language Models
by: Yadav, Sargam, et al.
Published: (2024)
by: Yadav, Sargam, et al.
Published: (2024)
Dealing with Annotator Disagreement in Hate Speech Classification
by: Dehghan, Somaiyeh, et al.
Published: (2025)
by: Dehghan, Somaiyeh, et al.
Published: (2025)
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech
by: Ludwig, Florian, et al.
Published: (2025)
by: Ludwig, Florian, et al.
Published: (2025)
Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages
by: Prome, Ruhina Tabasshum, et al.
Published: (2025)
by: Prome, Ruhina Tabasshum, et al.
Published: (2025)
A Federated Approach to Few-Shot Hate Speech Detection for Marginalized Communities
by: Ye, Haotian, et al.
Published: (2024)
by: Ye, Haotian, et al.
Published: (2024)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
by: Park, Junyeong, et al.
Published: (2025)
by: Park, Junyeong, et al.
Published: (2025)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
by: Almohaimeed, Saad, et al.
Published: (2025)
by: Almohaimeed, Saad, et al.
Published: (2025)
Hate Speech Detection using Large Language Models with Data Augmentation and Feature Enhancement
by: Nge, Brian Jing Hong, et al.
Published: (2026)
by: Nge, Brian Jing Hong, et al.
Published: (2026)
Soft Token Attacks Cannot Reliably Audit Unlearning in Large Language Models
by: Chen, Haokun, et al.
Published: (2025)
by: Chen, Haokun, et al.
Published: (2025)
Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
by: Lu, Junyu, et al.
Published: (2025)
by: Lu, Junyu, et al.
Published: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
by: Pres, Itamar, et al.
Published: (2024)
by: Pres, Itamar, et al.
Published: (2024)
An Investigation Into Explainable Audio Hate Speech Detection
by: An, Jinmyeong, et al.
Published: (2024)
by: An, Jinmyeong, et al.
Published: (2024)
Hierarchical Sentiment Analysis Framework for Hate Speech Detection: Implementing Binary and Multiclass Classification Strategy
by: Naznin, Faria, et al.
Published: (2024)
by: Naznin, Faria, et al.
Published: (2024)
More Than Sum of Its Parts: Deciphering Intent Shifts in Multimodal Hate Speech Detection
by: Sun, Runze, et al.
Published: (2026)
by: Sun, Runze, et al.
Published: (2026)
Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
by: Gajewska, Ewelina, et al.
Published: (2026)
by: Gajewska, Ewelina, et al.
Published: (2026)
Similar Items
-
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2024) -
Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models
by: Piot, Paloma, et al.
Published: (2025) -
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
by: Piot, Paloma, et al.
Published: (2024) -
Towards Efficient and Explainable Hate Speech Detection via Model Distillation
by: Piot, Paloma, et al.
Published: (2024) -
WATCHED: A Web AI Agent Tool for Combating Hate Speech by Expanding Data
by: Piot, Paloma, et al.
Published: (2025)