Disentangling Hate Across Target Identities
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Yiping, Wanner, Leo, Koya, Aneesh Moideen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Weakly-Supervised Hate Speech Classification Across Datasets
by: Jin, Yiping, et al.
Published: (2023)
by: Jin, Yiping, et al.
Published: (2023)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
by: Jin, Yiping, et al.
Published: (2024)
by: Jin, Yiping, et al.
Published: (2024)
Reasoning Models Will Sometimes Lie About Their Reasoning
by: Walden, William, et al.
Published: (2026)
by: Walden, William, et al.
Published: (2026)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
by: Proskurina, Irina, et al.
Published: (2025)
by: Proskurina, Irina, et al.
Published: (2025)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
by: Hu, Yujia, et al.
Published: (2026)
by: Hu, Yujia, et al.
Published: (2026)
xList-Hate: A Checklist-Based Framework for Interpretable and Generalizable Hate Speech Detection
by: Girón, Adrián, et al.
Published: (2026)
by: Girón, Adrián, et al.
Published: (2026)
Bryndza at ClimateActivism 2024: Stance, Target and Hate Event Detection via Retrieval-Augmented GPT-4 and LLaMA
by: Šuppa, Marek, et al.
Published: (2024)
by: Šuppa, Marek, et al.
Published: (2024)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
by: Kumarage, Tharindu, et al.
Published: (2024)
by: Kumarage, Tharindu, et al.
Published: (2024)
BengaliSent140: A Large-Scale Bengali Binary Sentiment Dataset for Hate and Non-Hate Speech Classification
by: Islam, Akif, et al.
Published: (2026)
by: Islam, Akif, et al.
Published: (2026)
Empirical Evaluation of Public HateSpeech Datasets
by: Jaf, Sadar, et al.
Published: (2024)
by: Jaf, Sadar, et al.
Published: (2024)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
by: Giorgi, Tommaso, et al.
Published: (2024)
by: Giorgi, Tommaso, et al.
Published: (2024)
Interpretable Discriminative Text Representations via Agreement and Label Disentanglement
by: Wang, Tong, et al.
Published: (2026)
by: Wang, Tong, et al.
Published: (2026)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
by: Lee, Nayeon, et al.
Published: (2023)
by: Lee, Nayeon, et al.
Published: (2023)
Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
by: Saha, Sougata, et al.
Published: (2024)
by: Saha, Sougata, et al.
Published: (2024)
The Impact of Persona-based Political Perspectives on Hateful Content Detection
by: Civelli, Stefano, et al.
Published: (2025)
by: Civelli, Stefano, et al.
Published: (2025)
Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection
by: Kim, Yumin, et al.
Published: (2025)
by: Kim, Yumin, et al.
Published: (2025)
ShED-HD: A Shannon Entropy Distribution Framework for Lightweight Hallucination Detection on Edge Devices
by: Vathul, Aneesh, et al.
Published: (2025)
by: Vathul, Aneesh, et al.
Published: (2025)
Who Speaks Matters: Analysing the Influence of the Speaker's Ethnicity on Hate Classification
by: Malik, Ananya, et al.
Published: (2024)
by: Malik, Ananya, et al.
Published: (2024)
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025)
by: Bauer, Julie, et al.
Published: (2025)
Disentangling the Roles of Target-Side Transfer and Regularization in Multilingual Machine Translation
by: Meng, Yan, et al.
Published: (2024)
by: Meng, Yan, et al.
Published: (2024)
1-800-SHARED-TASKS @ NLU of Devanagari Script Languages: Detection of Language, Hate Speech, and Targets using LLMs
by: Purbey, Jebish, et al.
Published: (2024)
by: Purbey, Jebish, et al.
Published: (2024)
A Federated Approach to Few-Shot Hate Speech Detection for Marginalized Communities
by: Ye, Haotian, et al.
Published: (2024)
by: Ye, Haotian, et al.
Published: (2024)
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech
by: Ludwig, Florian, et al.
Published: (2025)
by: Ludwig, Florian, et al.
Published: (2025)
Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages
by: Prome, Ruhina Tabasshum, et al.
Published: (2025)
by: Prome, Ruhina Tabasshum, et al.
Published: (2025)
Efficient Models for the Detection of Hate, Abuse and Profanity
by: Tillmann, Christoph, et al.
Published: (2024)
by: Tillmann, Christoph, et al.
Published: (2024)
GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Models
by: Hutson, Dylan, et al.
Published: (2025)
by: Hutson, Dylan, et al.
Published: (2025)
Hatred Stems from Ignorance! Distillation of the Persuasion Modes in Countering Conversational Hate Speech
by: Alyahya, Ghadi, et al.
Published: (2024)
by: Alyahya, Ghadi, et al.
Published: (2024)
Hate Speech Detection using Large Language Models with Data Augmentation and Feature Enhancement
by: Nge, Brian Jing Hong, et al.
Published: (2026)
by: Nge, Brian Jing Hong, et al.
Published: (2026)
CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Models
by: Komanduri, Aneesh, et al.
Published: (2025)
by: Komanduri, Aneesh, et al.
Published: (2025)
Target-driven Attack for Large Language Models
by: Zhang, Chong, et al.
Published: (2024)
by: Zhang, Chong, et al.
Published: (2024)
Hierarchical Sentiment Analysis Framework for Hate Speech Detection: Implementing Binary and Multiclass Classification Strategy
by: Naznin, Faria, et al.
Published: (2024)
by: Naznin, Faria, et al.
Published: (2024)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
More Than Sum of Its Parts: Deciphering Intent Shifts in Multimodal Hate Speech Detection
by: Sun, Runze, et al.
Published: (2026)
by: Sun, Runze, et al.
Published: (2026)
Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
by: Gajewska, Ewelina, et al.
Published: (2026)
by: Gajewska, Ewelina, et al.
Published: (2026)
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia
by: Ng, Ri Chi, et al.
Published: (2026)
by: Ng, Ri Chi, et al.
Published: (2026)
IndoToxic2024: A Demographically-Enriched Dataset of Hate Speech and Toxicity Types for Indonesian Language
by: Susanto, Lucky, et al.
Published: (2024)
by: Susanto, Lucky, et al.
Published: (2024)
A Survey of Machine Learning Models and Datasets for the Multi-label Classification of Textual Hate Speech in English
by: Bäumler, Julian, et al.
Published: (2025)
by: Bäumler, Julian, et al.
Published: (2025)
ToxSyn: Reducing Bias in Hate Speech Detection via Synthetic Minority Data in Brazilian Portuguese
by: Brito, Iago Alves, et al.
Published: (2025)
by: Brito, Iago Alves, et al.
Published: (2025)
Causal Intersectionality and Dual Form of Gradient Descent for Multimodal Analysis: a Case Study on Hateful Memes
by: Miyanishi, Yosuke, et al.
Published: (2023)
by: Miyanishi, Yosuke, et al.
Published: (2023)
Similar Items
-
Towards Weakly-Supervised Hate Speech Classification Across Datasets
by: Jin, Yiping, et al.
Published: (2023) -
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
by: Jin, Yiping, et al.
Published: (2024) -
Reasoning Models Will Sometimes Lie About Their Reasoning
by: Walden, William, et al.
Published: (2026) -
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
by: Proskurina, Irina, et al.
Published: (2025) -
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
by: Hu, Yujia, et al.
Published: (2026)