Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | Muscato, Benedetta, Chen, Beiduo, Gezici, Gizem, Plank, Barbara, Giannotti, Fosca |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Perspectives in Play: A Multi-Perspective Approach for More Inclusive NLP Systems
por: Muscato, Benedetta, et al.
Publicado: (2025)
por: Muscato, Benedetta, et al.
Publicado: (2025)
Multi-Perspective Stance Detection
por: Muscato, Benedetta, et al.
Publicado: (2024)
por: Muscato, Benedetta, et al.
Publicado: (2024)
Bridging the Gap: In-Context Learning for Modeling Human Disagreement
por: Muscato, Benedetta, et al.
Publicado: (2025)
por: Muscato, Benedetta, et al.
Publicado: (2025)
Embracing Diversity: A Multi-Perspective Approach with Soft Labels
por: Muscato, Benedetta, et al.
Publicado: (2025)
por: Muscato, Benedetta, et al.
Publicado: (2025)
Hybrid Retrieval for Hallucination Mitigation in Large Language Models: A Comparative Analysis
por: Mala, Chandana Sree, et al.
Publicado: (2025)
por: Mala, Chandana Sree, et al.
Publicado: (2025)
Learning by Surprise: Surplexity for Mitigating Model Collapse in Generative AI
por: Gambetta, Daniele, et al.
Publicado: (2024)
por: Gambetta, Daniele, et al.
Publicado: (2024)
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
por: Hong, Pingjun, et al.
Publicado: (2025)
por: Hong, Pingjun, et al.
Publicado: (2025)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
por: Zhao, Raoyuan, et al.
Publicado: (2025)
por: Zhao, Raoyuan, et al.
Publicado: (2025)
Diagnosing Hate Speech Classification: Where Do Humans and Machines Disagree, and Why?
por: Yang, Xilin
Publicado: (2024)
por: Yang, Xilin
Publicado: (2024)
Self-Explaining Hate Speech Detection with Moral Rationales
por: Vargas, Francielle, et al.
Publicado: (2026)
por: Vargas, Francielle, et al.
Publicado: (2026)
Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models
por: Cheng, Xinyuan, et al.
Publicado: (2026)
por: Cheng, Xinyuan, et al.
Publicado: (2026)
A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI
por: Chen, Beiduo, et al.
Publicado: (2024)
por: Chen, Beiduo, et al.
Publicado: (2024)
Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation
por: Chen, Beiduo, et al.
Publicado: (2025)
por: Chen, Beiduo, et al.
Publicado: (2025)
Compositional Generalisation for Explainable Hate Speech Detection
por: Calabrese, Agostina, et al.
Publicado: (2025)
por: Calabrese, Agostina, et al.
Publicado: (2025)
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
por: Eilertsen, Brage, et al.
Publicado: (2025)
por: Eilertsen, Brage, et al.
Publicado: (2025)
Legal Experts Disagree With Rationale Extraction Techniques for Explaining ECtHR Case Outcome Classification
por: Namazov, Mahammad, et al.
Publicado: (2026)
por: Namazov, Mahammad, et al.
Publicado: (2026)
Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective
por: Chen, Beiduo, et al.
Publicado: (2026)
por: Chen, Beiduo, et al.
Publicado: (2026)
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
por: Chen, Beiduo, et al.
Publicado: (2026)
por: Chen, Beiduo, et al.
Publicado: (2026)
LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference
por: Hong, Pingjun, et al.
Publicado: (2025)
por: Hong, Pingjun, et al.
Publicado: (2025)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
por: Chen, Beiduo, et al.
Publicado: (2024)
por: Chen, Beiduo, et al.
Publicado: (2024)
From Dissonance to Insights: Dissecting Disagreements in Rationale Construction for Case Outcome Classification
por: Xu, Shanshan, et al.
Publicado: (2023)
por: Xu, Shanshan, et al.
Publicado: (2023)
"Is Hate Lost in Translation?": Evaluation of Multilingual LGBTQIA+ Hate Speech Detection
por: Chan, Fai Leui, et al.
Publicado: (2024)
por: Chan, Fai Leui, et al.
Publicado: (2024)
Towards Efficient and Explainable Hate Speech Detection via Model Distillation
por: Piot, Paloma, et al.
Publicado: (2024)
por: Piot, Paloma, et al.
Publicado: (2024)
An Investigation Into Explainable Audio Hate Speech Detection
por: An, Jinmyeong, et al.
Publicado: (2024)
por: An, Jinmyeong, et al.
Publicado: (2024)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
por: Nirmal, Ayushi, et al.
Publicado: (2024)
por: Nirmal, Ayushi, et al.
Publicado: (2024)
NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative Data
por: Tonneau, Manuel, et al.
Publicado: (2024)
por: Tonneau, Manuel, et al.
Publicado: (2024)
Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
por: Singh, Daman Deep, et al.
Publicado: (2025)
por: Singh, Daman Deep, et al.
Publicado: (2025)
Hate Speech Detection and Classification in Amharic Text with Deep Learning
por: Gashe, Samuel Minale, et al.
Publicado: (2024)
por: Gashe, Samuel Minale, et al.
Publicado: (2024)
Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
por: Blaschke, Verena, et al.
Publicado: (2025)
por: Blaschke, Verena, et al.
Publicado: (2025)
Cracking the Code: Enhancing Implicit Hate Speech Detection through Coding Classification
por: Wei, Lu, et al.
Publicado: (2025)
por: Wei, Lu, et al.
Publicado: (2025)
Rethinking Human Preference Evaluation of LLM Rationales
por: Li, Ziang, et al.
Publicado: (2025)
por: Li, Ziang, et al.
Publicado: (2025)
Incorporating Human Explanations for Robust Hate Speech Detection
por: Chen, Jennifer L., et al.
Publicado: (2024)
por: Chen, Jennifer L., et al.
Publicado: (2024)
MasonPerplexity at Multimodal Hate Speech Event Detection 2024: Hate Speech and Target Detection Using Transformer Ensembles
por: Ganguly, Amrita, et al.
Publicado: (2024)
por: Ganguly, Amrita, et al.
Publicado: (2024)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
por: Yuan, Shuzhou, et al.
Publicado: (2025)
por: Yuan, Shuzhou, et al.
Publicado: (2025)
HateDebias: On the Diversity and Variability of Hate Speech Debiasing
por: Wu, Hongyan, et al.
Publicado: (2024)
por: Wu, Hongyan, et al.
Publicado: (2024)
Automatic Textual Normalization for Hate Speech Detection
por: Nguyen, Anh Thi-Hoang, et al.
Publicado: (2023)
por: Nguyen, Anh Thi-Hoang, et al.
Publicado: (2023)
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
por: Wang, Yifan, et al.
Publicado: (2025)
por: Wang, Yifan, et al.
Publicado: (2025)
Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster
por: Calabrese, Agostina, et al.
Publicado: (2024)
por: Calabrese, Agostina, et al.
Publicado: (2024)
Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection
por: Wang, Bing, et al.
Publicado: (2026)
por: Wang, Bing, et al.
Publicado: (2026)
Demystifying Hateful Content: Leveraging Large Multimodal Models for Hateful Meme Detection with Explainable Decisions
por: Hee, Ming Shan, et al.
Publicado: (2025)
por: Hee, Ming Shan, et al.
Publicado: (2025)
Ejemplares similares
-
Perspectives in Play: A Multi-Perspective Approach for More Inclusive NLP Systems
por: Muscato, Benedetta, et al.
Publicado: (2025) -
Multi-Perspective Stance Detection
por: Muscato, Benedetta, et al.
Publicado: (2024) -
Bridging the Gap: In-Context Learning for Modeling Human Disagreement
por: Muscato, Benedetta, et al.
Publicado: (2025) -
Embracing Diversity: A Multi-Perspective Approach with Soft Labels
por: Muscato, Benedetta, et al.
Publicado: (2025) -
Hybrid Retrieval for Hallucination Mitigation in Large Language Models: A Comparative Analysis
por: Mala, Chandana Sree, et al.
Publicado: (2025)