HateDebias: On the Diversity and Variability of Hate Speech Debiasing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Hongyan, Chen, Zhengming, Li, Zijian, Lin, Nankai, Wang, Lianxi, Jiang, Shengyi, Yang, Aimin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLASS: Enhancing Cross-Modal Text-Molecule Retrieval Performance and Training Efficiency
von: Wu, Hongyan, et al.
Veröffentlicht: (2025)
von: Wu, Hongyan, et al.
Veröffentlicht: (2025)
A Simple Yet Effective Corpus Construction Framework for Indonesian Grammatical Error Correction
von: Lin, Nankai, et al.
Veröffentlicht: (2024)
von: Lin, Nankai, et al.
Veröffentlicht: (2024)
Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
Web(er) of Hate: A Survey on How Hate Speech Is Typed
von: Wang, Luna, et al.
Veröffentlicht: (2025)
von: Wang, Luna, et al.
Veröffentlicht: (2025)
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
Multiple-Debias: A Full-process Debiasing Method for Multilingual Pre-trained Language Models
von: Liang, Haoyu, et al.
Veröffentlicht: (2026)
von: Liang, Haoyu, et al.
Veröffentlicht: (2026)
Evaluating Simple Debiasing Techniques in RoBERTa-based Hate Speech Detection Models
von: Iftimie, Diana, et al.
Veröffentlicht: (2025)
von: Iftimie, Diana, et al.
Veröffentlicht: (2025)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
"Is Hate Lost in Translation?": Evaluation of Multilingual LGBTQIA+ Hate Speech Detection
von: Chan, Fai Leui, et al.
Veröffentlicht: (2024)
von: Chan, Fai Leui, et al.
Veröffentlicht: (2024)
Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
HateModerate: Testing Hate Speech Detectors against Content Moderation Policies
von: Zheng, Jiangrui, et al.
Veröffentlicht: (2023)
von: Zheng, Jiangrui, et al.
Veröffentlicht: (2023)
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
von: Chapagain, Santosh, et al.
Veröffentlicht: (2025)
von: Chapagain, Santosh, et al.
Veröffentlicht: (2025)
HateGPT: Unleashing GPT-3.5 Turbo to Combat Hate Speech on X
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative Data
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
Incorporating Human Explanations for Robust Hate Speech Detection
von: Chen, Jennifer L., et al.
Veröffentlicht: (2024)
von: Chen, Jennifer L., et al.
Veröffentlicht: (2024)
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
von: Proskurina, Irina, et al.
Veröffentlicht: (2025)
von: Proskurina, Irina, et al.
Veröffentlicht: (2025)
EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models
von: Bui, Minh Duc, et al.
Veröffentlicht: (2024)
von: Bui, Minh Duc, et al.
Veröffentlicht: (2024)
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages
von: Muhammad, Shamsuddeen Hassan, et al.
Veröffentlicht: (2025)
von: Muhammad, Shamsuddeen Hassan, et al.
Veröffentlicht: (2025)
Dialogues of Dissent: Thematic and Rhetorical Dimensions of Hate and Counter-Hate Speech in Social Media Conversations
von: Levi, Effi, et al.
Veröffentlicht: (2025)
von: Levi, Effi, et al.
Veröffentlicht: (2025)
HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation
von: Rizwan, Naquee, et al.
Veröffentlicht: (2025)
von: Rizwan, Naquee, et al.
Veröffentlicht: (2025)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
von: Jin, Yiping, et al.
Veröffentlicht: (2024)
von: Jin, Yiping, et al.
Veröffentlicht: (2024)
MasonPerplexity at Multimodal Hate Speech Event Detection 2024: Hate Speech and Target Detection Using Transformer Ensembles
von: Ganguly, Amrita, et al.
Veröffentlicht: (2024)
von: Ganguly, Amrita, et al.
Veröffentlicht: (2024)
Few-shot Hate Speech Detection Based on the MindSpore Framework
von: Qin, Zhenkai, et al.
Veröffentlicht: (2025)
von: Qin, Zhenkai, et al.
Veröffentlicht: (2025)
HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
von: Sen, Tanmay, et al.
Veröffentlicht: (2024)
von: Sen, Tanmay, et al.
Veröffentlicht: (2024)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
von: Hu, Yujia, et al.
Veröffentlicht: (2026)
von: Hu, Yujia, et al.
Veröffentlicht: (2026)
Deciphering Hate: Identifying Hateful Memes and Their Targets
von: Hossain, Eftekhar, et al.
Veröffentlicht: (2024)
von: Hossain, Eftekhar, et al.
Veröffentlicht: (2024)
EkoHate: Abusive Language and Hate Speech Detection for Code-switched Political Discussions on Nigerian Twitter
von: Ilevbare, Comfort Eseohen, et al.
Veröffentlicht: (2024)
von: Ilevbare, Comfort Eseohen, et al.
Veröffentlicht: (2024)
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
Compositional Generalisation for Explainable Hate Speech Detection
von: Calabrese, Agostina, et al.
Veröffentlicht: (2025)
von: Calabrese, Agostina, et al.
Veröffentlicht: (2025)
Automatic Textual Normalization for Hate Speech Detection
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2023)
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2023)
Challenger at MultiPRIDE: Is It Hate Speech or Reclaimed?
von: Tekanlou, Hadi Bayrami Asl, et al.
Veröffentlicht: (2026)
von: Tekanlou, Hadi Bayrami Asl, et al.
Veröffentlicht: (2026)
The Enforcement and Feasibility of Hate Speech Moderation on Twitter
von: Tonneau, Manuel, et al.
Veröffentlicht: (2026)
von: Tonneau, Manuel, et al.
Veröffentlicht: (2026)
Jailbreaking? One Step Is Enough!
von: Zheng, Weixiong, et al.
Veröffentlicht: (2024)
von: Zheng, Weixiong, et al.
Veröffentlicht: (2024)
Hate Speech Detection with Generalizable Target-aware Fairness
von: Chen, Tong, et al.
Veröffentlicht: (2024)
von: Chen, Tong, et al.
Veröffentlicht: (2024)
ReZG: Retrieval-Augmented Zero-Shot Counter Narrative Generation for Hate Speech
von: Jiang, Shuyu, et al.
Veröffentlicht: (2023)
von: Jiang, Shuyu, et al.
Veröffentlicht: (2023)
DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images
von: Dalal, Dwip, et al.
Veröffentlicht: (2025)
von: Dalal, Dwip, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLASS: Enhancing Cross-Modal Text-Molecule Retrieval Performance and Training Efficiency
von: Wu, Hongyan, et al.
Veröffentlicht: (2025) -
A Simple Yet Effective Corpus Construction Framework for Indonesian Grammatical Error Correction
von: Lin, Nankai, et al.
Veröffentlicht: (2024) -
Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models
von: Piot, Paloma, et al.
Veröffentlicht: (2025) -
Web(er) of Hate: A Survey on How Hate Speech Is Typed
von: Wang, Luna, et al.
Veröffentlicht: (2025) -
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
von: Piot, Paloma, et al.
Veröffentlicht: (2024)