Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Kumarage, Tharindu, Bhattacharjee, Amrita, Garland, Joshua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
by: Nirmal, Ayushi, et al.
Published: (2024)
by: Nirmal, Ayushi, et al.
Published: (2024)
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech
by: Ludwig, Florian, et al.
Published: (2025)
by: Ludwig, Florian, et al.
Published: (2025)
Cross-Platform Hate Speech Detection with Weakly Supervised Causal Disentanglement
by: Sheth, Paras, et al.
Published: (2024)
by: Sheth, Paras, et al.
Published: (2024)
Efficient Models for the Detection of Hate, Abuse and Profanity
by: Tillmann, Christoph, et al.
Published: (2024)
by: Tillmann, Christoph, et al.
Published: (2024)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
by: Proskurina, Irina, et al.
Published: (2025)
by: Proskurina, Irina, et al.
Published: (2025)
A Survey of AI-generated Text Forensic Systems: Detection, Attribution, and Characterization
by: Kumarage, Tharindu, et al.
Published: (2024)
by: Kumarage, Tharindu, et al.
Published: (2024)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
by: Das, Amit, et al.
Published: (2024)
by: Das, Amit, et al.
Published: (2024)
Hate Speech Detection using Large Language Models with Data Augmentation and Feature Enhancement
by: Nge, Brian Jing Hong, et al.
Published: (2026)
by: Nge, Brian Jing Hong, et al.
Published: (2026)
EAGLE: A Domain Generalization Framework for AI-generated Text Detection
by: Bhattacharjee, Amrita, et al.
Published: (2024)
by: Bhattacharjee, Amrita, et al.
Published: (2024)
xList-Hate: A Checklist-Based Framework for Interpretable and Generalizable Hate Speech Detection
by: Girón, Adrián, et al.
Published: (2026)
by: Girón, Adrián, et al.
Published: (2026)
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
by: Bhattacharjee, Amrita, et al.
Published: (2024)
by: Bhattacharjee, Amrita, et al.
Published: (2024)
An Investigation of Large Language Models for Real-World Hate Speech Detection
by: Guo, Keyan, et al.
Published: (2024)
by: Guo, Keyan, et al.
Published: (2024)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
by: Hu, Yujia, et al.
Published: (2026)
by: Hu, Yujia, et al.
Published: (2026)
Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages
by: Prome, Ruhina Tabasshum, et al.
Published: (2025)
by: Prome, Ruhina Tabasshum, et al.
Published: (2025)
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025)
by: Bauer, Julie, et al.
Published: (2025)
BengaliSent140: A Large-Scale Bengali Binary Sentiment Dataset for Hate and Non-Hate Speech Classification
by: Islam, Akif, et al.
Published: (2026)
by: Islam, Akif, et al.
Published: (2026)
Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection
by: Kim, Yumin, et al.
Published: (2025)
by: Kim, Yumin, et al.
Published: (2025)
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
by: Bhattacharjee, Amrita, et al.
Published: (2023)
by: Bhattacharjee, Amrita, et al.
Published: (2023)
Multilingual Hate Speech Detection in Social Media Using Translation-Based Approaches with Large Language Models
by: Usman, Muhammad, et al.
Published: (2025)
by: Usman, Muhammad, et al.
Published: (2025)
Empirical Evaluation of Public HateSpeech Datasets
by: Jaf, Sadar, et al.
Published: (2024)
by: Jaf, Sadar, et al.
Published: (2024)
Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI
by: Böck, Adrian Jaques, et al.
Published: (2024)
by: Böck, Adrian Jaques, et al.
Published: (2024)
OSPC: Artificial VLM Features for Hateful Meme Detection
by: Grönquist, Peter
Published: (2024)
by: Grönquist, Peter
Published: (2024)
An Investigation Into Explainable Audio Hate Speech Detection
by: An, Jinmyeong, et al.
Published: (2024)
by: An, Jinmyeong, et al.
Published: (2024)
SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia
by: Ng, Ri Chi, et al.
Published: (2026)
by: Ng, Ri Chi, et al.
Published: (2026)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
by: Nghiem, Huy, et al.
Published: (2024)
by: Nghiem, Huy, et al.
Published: (2024)
A Federated Approach to Few-Shot Hate Speech Detection for Marginalized Communities
by: Ye, Haotian, et al.
Published: (2024)
by: Ye, Haotian, et al.
Published: (2024)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
by: Almohaimeed, Saad, et al.
Published: (2025)
by: Almohaimeed, Saad, et al.
Published: (2025)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
by: Lee, Nayeon, et al.
Published: (2023)
by: Lee, Nayeon, et al.
Published: (2023)
Feature Selection Empowered BERT for Detection of Hate Speech with Vocabulary Augmentation
by: Desai, Pritish N., et al.
Published: (2025)
by: Desai, Pritish N., et al.
Published: (2025)
Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
by: Saha, Sougata, et al.
Published: (2024)
by: Saha, Sougata, et al.
Published: (2024)
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
by: Piot, Paloma, et al.
Published: (2024)
by: Piot, Paloma, et al.
Published: (2024)
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
by: Zhao, Chengshuai, et al.
Published: (2025)
by: Zhao, Chengshuai, et al.
Published: (2025)
Leveraging Weakly Annotated Data for Hate Speech Detection in Code-Mixed Hinglish: A Feasibility-Driven Transfer Learning Approach with Large Language Models
by: Yadav, Sargam, et al.
Published: (2024)
by: Yadav, Sargam, et al.
Published: (2024)
System Report for CCL25-Eval Task 10: Prompt-Driven Large Language Model Merge for Fine-Grained Chinese Hate Speech Detection
by: Wu, Binglin, et al.
Published: (2025)
by: Wu, Binglin, et al.
Published: (2025)
Seeing Hate Differently: Hate Subspace Modeling for Culture-Aware Hate Speech Detection
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
Hierarchical Sentiment Analysis Framework for Hate Speech Detection: Implementing Binary and Multiclass Classification Strategy
by: Naznin, Faria, et al.
Published: (2024)
by: Naznin, Faria, et al.
Published: (2024)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Similar Items
-
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
by: Nirmal, Ayushi, et al.
Published: (2024) -
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech
by: Ludwig, Florian, et al.
Published: (2025) -
Cross-Platform Hate Speech Detection with Weakly Supervised Causal Disentanglement
by: Sheth, Paras, et al.
Published: (2024) -
Efficient Models for the Detection of Hate, Abuse and Profanity
by: Tillmann, Christoph, et al.
Published: (2024) -
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
by: Proskurina, Irina, et al.
Published: (2025)