Towards Weakly-Supervised Hate Speech Classification Across Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Yiping, Wanner, Leo, Kadam, Vishakha Laxman, Shvets, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
by: Jin, Yiping, et al.
Published: (2024)
by: Jin, Yiping, et al.
Published: (2024)
Disentangling Hate Across Target Identities
by: Jin, Yiping, et al.
Published: (2024)
by: Jin, Yiping, et al.
Published: (2024)
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025)
by: Bauer, Julie, et al.
Published: (2025)
Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
BengaliSent140: A Large-Scale Bengali Binary Sentiment Dataset for Hate and Non-Hate Speech Classification
by: Islam, Akif, et al.
Published: (2026)
by: Islam, Akif, et al.
Published: (2026)
An Investigation of Large Language Models for Real-World Hate Speech Detection
by: Guo, Keyan, et al.
Published: (2024)
by: Guo, Keyan, et al.
Published: (2024)
Empirical Evaluation of Public HateSpeech Datasets
by: Jaf, Sadar, et al.
Published: (2024)
by: Jaf, Sadar, et al.
Published: (2024)
Subjective $\textit{Isms}$? On the Danger of Conflating Hate and Offence in Abusive Language Detection
by: Curry, Amanda Cercas, et al.
Published: (2024)
by: Curry, Amanda Cercas, et al.
Published: (2024)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
by: Almohaimeed, Saad, et al.
Published: (2025)
by: Almohaimeed, Saad, et al.
Published: (2025)
MemeScouts@LT-EDI 2026: Asking the Right Questions -- Prompted Weak Supervision for Meme Hate Speech Detection
by: Bueno, Ivo, et al.
Published: (2026)
by: Bueno, Ivo, et al.
Published: (2026)
A Survey of Machine Learning Models and Datasets for the Multi-label Classification of Textual Hate Speech in English
by: Bäumler, Julian, et al.
Published: (2025)
by: Bäumler, Julian, et al.
Published: (2025)
DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
by: Fu, Jiachen, et al.
Published: (2025)
by: Fu, Jiachen, et al.
Published: (2025)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
by: Lee, Nayeon, et al.
Published: (2023)
by: Lee, Nayeon, et al.
Published: (2023)
Bridging the Data Provenance Gap Across Text, Speech and Video
by: Longpre, Shayne, et al.
Published: (2024)
by: Longpre, Shayne, et al.
Published: (2024)
Improving Hate Speech Classification with Cross-Taxonomy Dataset Integration
by: Fillies, Jan, et al.
Published: (2025)
by: Fillies, Jan, et al.
Published: (2025)
AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
by: Lee, Yejin, et al.
Published: (2025)
by: Lee, Yejin, et al.
Published: (2025)
Anecdoctoring: Automated Red-Teaming Across Language and Place
by: Cuevas, Alejandro, et al.
Published: (2025)
by: Cuevas, Alejandro, et al.
Published: (2025)
Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
Towards Best Practices for Open Datasets for LLM Training
by: Baack, Stefan, et al.
Published: (2025)
by: Baack, Stefan, et al.
Published: (2025)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
by: Proskurina, Irina, et al.
Published: (2025)
by: Proskurina, Irina, et al.
Published: (2025)
NoisyHate: Mining Online Human-Written Perturbations for Realistic Robustness Benchmarking of Content Moderation Models
by: Ye, Yiran, et al.
Published: (2023)
by: Ye, Yiran, et al.
Published: (2023)
The Algorithmic Caricature: Auditing LLM-Generated Political Discourse Across Crisis Events
by: Gunjan, et al.
Published: (2026)
by: Gunjan, et al.
Published: (2026)
ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models
by: Heakl, Ahmed, et al.
Published: (2024)
by: Heakl, Ahmed, et al.
Published: (2024)
Hierarchical Sentiment Analysis Framework for Hate Speech Detection: Implementing Binary and Multiclass Classification Strategy
by: Naznin, Faria, et al.
Published: (2024)
by: Naznin, Faria, et al.
Published: (2024)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
by: Ding, Junchen, et al.
Published: (2025)
by: Ding, Junchen, et al.
Published: (2025)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
by: Hu, Yujia, et al.
Published: (2026)
by: Hu, Yujia, et al.
Published: (2026)
"Sorry, I Didn't Catch That": How Speech Models Miss What Matters Most
by: Zhou, Kaitlyn, et al.
Published: (2026)
by: Zhou, Kaitlyn, et al.
Published: (2026)
How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models
by: Fukui, Hiroki
Published: (2026)
by: Fukui, Hiroki
Published: (2026)
IndoToxic2024: A Demographically-Enriched Dataset of Hate Speech and Toxicity Types for Indonesian Language
by: Susanto, Lucky, et al.
Published: (2024)
by: Susanto, Lucky, et al.
Published: (2024)
Efficient Multilingual Name Type Classification Using Convolutional Networks
by: Lauc, Davor
Published: (2026)
by: Lauc, Davor
Published: (2026)
xList-Hate: A Checklist-Based Framework for Interpretable and Generalizable Hate Speech Detection
by: Girón, Adrián, et al.
Published: (2026)
by: Girón, Adrián, et al.
Published: (2026)
From Black-Box Confidence to Measurable Trust in Clinical AI: A Framework for Evidence, Supervision, and Staged Autonomy
by: Zabolotnii, Serhii, et al.
Published: (2026)
by: Zabolotnii, Serhii, et al.
Published: (2026)
SynBullying: A Multi LLM Synthetic Conversational Dataset for Cyberbullying Detection
by: Kazemi, Arefeh, et al.
Published: (2025)
by: Kazemi, Arefeh, et al.
Published: (2025)
Towards Safe Multilingual Frontier AI
by: Kanepajs, Artūrs, et al.
Published: (2024)
by: Kanepajs, Artūrs, et al.
Published: (2024)
Understanding Social Support Needs in Questions: A Hybrid Approach Integrating Semi-Supervised Learning and LLM-based Data Augmentation
by: Kuang, Junwei, et al.
Published: (2025)
by: Kuang, Junwei, et al.
Published: (2025)
That's So FETCH: Fashioning Ensemble Techniques for LLM Classification in Civil Legal Intake and Referral
by: Steenhuis, Quinten
Published: (2025)
by: Steenhuis, Quinten
Published: (2025)
DeepTutor: Towards Agentic Personalized Tutoring
by: Zhao, Bingxi, et al.
Published: (2026)
by: Zhao, Bingxi, et al.
Published: (2026)
Towards Next-Generation Medical Agent: How o1 is Reshaping Decision-Making in Medical Scenarios
by: Xu, Shaochen, et al.
Published: (2024)
by: Xu, Shaochen, et al.
Published: (2024)
The Unseen Targets of Hate -- A Systematic Review of Hateful Communication Datasets
by: Yu, Zehui, et al.
Published: (2024)
by: Yu, Zehui, et al.
Published: (2024)
Towards medical AI misalignment: a preliminary study
by: Puccio, Barbara, et al.
Published: (2025)
by: Puccio, Barbara, et al.
Published: (2025)
Similar Items
-
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
by: Jin, Yiping, et al.
Published: (2024) -
Disentangling Hate Across Target Identities
by: Jin, Yiping, et al.
Published: (2024) -
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025) -
Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection
by: Ahmed, Ahmed Haj, et al.
Published: (2024) -
BengaliSent140: A Large-Scale Bengali Binary Sentiment Dataset for Hate and Non-Hate Speech Classification
by: Islam, Akif, et al.
Published: (2026)