Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Eilertsen, Brage, Bjørgfinsdóttir, Røskva, Vargas, Francielle, Ramezani-Kebrya, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Explaining Hate Speech Detection with Moral Rationales
by: Vargas, Francielle, et al.
Published: (2026)
by: Vargas, Francielle, et al.
Published: (2026)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
by: Nirmal, Ayushi, et al.
Published: (2024)
by: Nirmal, Ayushi, et al.
Published: (2024)
Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
by: Chapagain, Santosh, et al.
Published: (2025)
by: Chapagain, Santosh, et al.
Published: (2025)
HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
by: Sen, Tanmay, et al.
Published: (2024)
by: Sen, Tanmay, et al.
Published: (2024)
Hate Speech Detection with Generalizable Target-aware Fairness
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
by: Indrehus, Kjetil, et al.
Published: (2026)
by: Indrehus, Kjetil, et al.
Published: (2026)
Hate Speech Detection and Classification in Amharic Text with Deep Learning
by: Gashe, Samuel Minale, et al.
Published: (2024)
by: Gashe, Samuel Minale, et al.
Published: (2024)
An Effective, Robust and Fairness-aware Hate Speech Detection Framework
by: Mou, Guanyi, et al.
Published: (2024)
by: Mou, Guanyi, et al.
Published: (2024)
Cross-Platform Hate Speech Detection with Weakly Supervised Causal Disentanglement
by: Sheth, Paras, et al.
Published: (2024)
by: Sheth, Paras, et al.
Published: (2024)
A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages
by: Das, Susmita, et al.
Published: (2024)
by: Das, Susmita, et al.
Published: (2024)
Diagnosing Hate Speech Classification: Where Do Humans and Machines Disagree, and Why?
by: Yang, Xilin
Published: (2024)
by: Yang, Xilin
Published: (2024)
BOISHOMMO: Holistic Approach for Bangla Hate Speech
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
by: Almohaimeed, Saad, et al.
Published: (2025)
by: Almohaimeed, Saad, et al.
Published: (2025)
Deep Learning Approaches for Detecting Adversarial Cyberbullying and Hate Speech in Social Networks
by: Azumah, Sylvia Worlali, et al.
Published: (2024)
by: Azumah, Sylvia Worlali, et al.
Published: (2024)
Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs
by: Kaiser, Daniel, et al.
Published: (2026)
by: Kaiser, Daniel, et al.
Published: (2026)
On Importance of Code-Mixed Embeddings for Hate Speech Identification
by: Jagdale, Shruti, et al.
Published: (2024)
by: Jagdale, Shruti, et al.
Published: (2024)
SWE2: SubWord Enriched and Significant Word Emphasized Framework for Hate Speech Detection
by: Mou, Guanyi, et al.
Published: (2024)
by: Mou, Guanyi, et al.
Published: (2024)
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
by: Zhao, Chengshuai, et al.
Published: (2025)
by: Zhao, Chengshuai, et al.
Published: (2025)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
by: Das, Amit, et al.
Published: (2024)
by: Das, Amit, et al.
Published: (2024)
Factuality and Transparency Are All RAG Needs! Self-Explaining Contrastive Evidence Re-ranking
by: Vargas, Francielle, et al.
Published: (2025)
by: Vargas, Francielle, et al.
Published: (2025)
Transformers and Ensemble methods: A solution for Hate Speech Detection in Arabic languages
by: de Paula, Angel Felipe Magnossão, et al.
Published: (2023)
by: de Paula, Angel Felipe Magnossão, et al.
Published: (2023)
Comparison of Scoring Rationales Between Large Language Models and Human Raters
by: Hua, Haowei, et al.
Published: (2025)
by: Hua, Haowei, et al.
Published: (2025)
Principled Operator Learning in Ocean Dynamics: The Role of Temporal Structure
by: Jahanmard, Vahidreza, et al.
Published: (2025)
by: Jahanmard, Vahidreza, et al.
Published: (2025)
Aligning Human and Machine Attention for Enhanced Supervised Learning
by: Chriqui, Avihay, et al.
Published: (2025)
by: Chriqui, Avihay, et al.
Published: (2025)
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
by: Muscato, Benedetta, et al.
Published: (2026)
by: Muscato, Benedetta, et al.
Published: (2026)
From BERT to Qwen: Hate Detection across architectures
by: Mon, Ariadna, et al.
Published: (2025)
by: Mon, Ariadna, et al.
Published: (2025)
Transfer Learning via Lexical Relatedness: A Sarcasm and Hate Speech Case Study
by: Cabrera, Angelly, et al.
Published: (2025)
by: Cabrera, Angelly, et al.
Published: (2025)
InstructRAG: Instructing Retrieval-Augmented Generation via Self-Synthesized Rationales
by: Wei, Zhepei, et al.
Published: (2024)
by: Wei, Zhepei, et al.
Published: (2024)
Multilingual Hate Speech Detection in Social Media Using Translation-Based Approaches with Large Language Models
by: Usman, Muhammad, et al.
Published: (2025)
by: Usman, Muhammad, et al.
Published: (2025)
Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI
by: Böck, Adrian Jaques, et al.
Published: (2024)
by: Böck, Adrian Jaques, et al.
Published: (2024)
Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch
by: Shukla, Prarabdh, et al.
Published: (2025)
by: Shukla, Prarabdh, et al.
Published: (2025)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
by: Xu, Tianyang, et al.
Published: (2024)
by: Xu, Tianyang, et al.
Published: (2024)
An Investigation of Large Language Models for Real-World Hate Speech Detection
by: Guo, Keyan, et al.
Published: (2024)
by: Guo, Keyan, et al.
Published: (2024)
How Ambiguous Are the Rationales for Natural Language Reasoning? A Simple Approach to Handling Rationale Uncertainty
by: Kim, Hazel H.
Published: (2024)
by: Kim, Hazel H.
Published: (2024)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
by: Mihaila, George
Published: (2026)
by: Mihaila, George
Published: (2026)
Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
by: Islam, Akif, et al.
Published: (2025)
by: Islam, Akif, et al.
Published: (2025)
1-800-SHARED-TASKS @ NLU of Devanagari Script Languages: Detection of Language, Hate Speech, and Targets using LLMs
by: Purbey, Jebish, et al.
Published: (2024)
by: Purbey, Jebish, et al.
Published: (2024)
Comparison of Modern Multilingual Text Embedding Techniques for Hate Speech Detection Task
by: Vaiciukynas, Evaldas, et al.
Published: (2026)
by: Vaiciukynas, Evaldas, et al.
Published: (2026)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
by: Kaiser, Daniel, et al.
Published: (2025)
by: Kaiser, Daniel, et al.
Published: (2025)
Similar Items
-
Self-Explaining Hate Speech Detection with Moral Rationales
by: Vargas, Francielle, et al.
Published: (2026) -
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
by: Nirmal, Ayushi, et al.
Published: (2024) -
Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
by: Chapagain, Santosh, et al.
Published: (2025) -
HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
by: Sen, Tanmay, et al.
Published: (2024) -
Hate Speech Detection with Generalizable Target-aware Fairness
by: Chen, Tong, et al.
Published: (2024)