Self-Explaining Hate Speech Detection with Moral Rationales
Fuente:
arXiv
Saved in:
| Main Authors: | Vargas, Francielle, Trager, Jackson, Alves, Diego, Thapa, Surendrabikram, Guida, Matteo, Atil, Berk, Dementieva, Daryna, Smart, Andrew, Agrawal, Ameeta |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation
by: Trager, Jackson, et al.
Published: (2025)
by: Trager, Jackson, et al.
Published: (2025)
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
by: Eilertsen, Brage, et al.
Published: (2025)
by: Eilertsen, Brage, et al.
Published: (2025)
Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
Explainable Semantic Textual Similarity via Dissimilar Span Detection
by: Lozano, Diego Miguel, et al.
Published: (2026)
by: Lozano, Diego Miguel, et al.
Published: (2026)
EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainian
by: Dementieva, Daryna, et al.
Published: (2025)
by: Dementieva, Daryna, et al.
Published: (2025)
CrossNews-UA: A Cross-lingual News Semantic Similarity Benchmark for Ukrainian, Polish, Russian, and English
by: Dementieva, Daryna, et al.
Published: (2025)
by: Dementieva, Daryna, et al.
Published: (2025)
MultiParaDetox: Extending Text Detoxification with Parallel Data to New Languages
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
Cross-lingual Text Classification Transfer: The Case of Ukrainian
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
Enhancing ESG Impact Type Identification through Early Fusion and Multilingual Models
by: Veeramani, Hariram, et al.
Published: (2024)
by: Veeramani, Hariram, et al.
Published: (2024)
franciellevargas/HateBR: v7.0.0
by: Francielle Vargas, et al.
Published: (2025)
by: Francielle Vargas, et al.
Published: (2025)
Factuality and Transparency Are All RAG Needs! Self-Explaining Contrastive Evidence Re-ranking
by: Vargas, Francielle, et al.
Published: (2025)
by: Vargas, Francielle, et al.
Published: (2025)
HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation
by: Rizwan, Naquee, et al.
Published: (2025)
by: Rizwan, Naquee, et al.
Published: (2025)
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
by: Muscato, Benedetta, et al.
Published: (2026)
by: Muscato, Benedetta, et al.
Published: (2026)
Evaluating Text Style Transfer: A Nine-Language Benchmark for Text Detoxification
by: Protasov, Vitaly, et al.
Published: (2025)
by: Protasov, Vitaly, et al.
Published: (2025)
Toxicity Classification in Ukrainian
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
by: Atil, Berk, et al.
Published: (2026)
by: Atil, Berk, et al.
Published: (2026)
Beyond Topical Similarity: Contrastive Evidence Retrieval with Interpretable Attention Alignment in RAG
by: Vargas, Francielle, et al.
Published: (2026)
by: Vargas, Francielle, et al.
Published: (2026)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
by: Nirmal, Ayushi, et al.
Published: (2024)
by: Nirmal, Ayushi, et al.
Published: (2024)
Explain-then-Process: Using Grammar Prompting to Enhance Grammatical Acceptability Judgments
by: Scheinberg, Russell, et al.
Published: (2025)
by: Scheinberg, Russell, et al.
Published: (2025)
Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian
by: Üyük, Cem, et al.
Published: (2024)
by: Üyük, Cem, et al.
Published: (2024)
No-Worse Context-Aware Decoding: Preventing Neutral Regression in Context-Conditioned Generation
by: Tao, Yufei, et al.
Published: (2026)
by: Tao, Yufei, et al.
Published: (2026)
MTQ-Eval: Multilingual Text Quality Evaluation for Language Models
by: Pokharel, Rhitabrat, et al.
Published: (2025)
by: Pokharel, Rhitabrat, et al.
Published: (2025)
Understanding Position Bias Effects on Fairness in Social Multi-Document Summarization
by: Olabisi, Olubusayo, et al.
Published: (2024)
by: Olabisi, Olubusayo, et al.
Published: (2024)
Explain the Flag: Contextualizing Hate Speech Beyond Censorship
by: Liartis, Jason, et al.
Published: (2026)
by: Liartis, Jason, et al.
Published: (2026)
Model Unlearning Objectives Vary for Distinct Language Functions
by: Atil, Berk, et al.
Published: (2026)
by: Atil, Berk, et al.
Published: (2026)
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
Something Just Like TRuST : Toxicity Recognition of Span and Target
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
Temperature Matters: Enhancing Watermark Robustness Against Paraphrasing Attacks
by: Idrissi, Badr Youbi, et al.
Published: (2025)
by: Idrissi, Badr Youbi, et al.
Published: (2025)
Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization
by: Giabbanelli, Philippe J., et al.
Published: (2025)
by: Giabbanelli, Philippe J., et al.
Published: (2025)
What Drives Performance in Multilingual Language Models?
by: Nezhad, Sina Bagheri, et al.
Published: (2024)
by: Nezhad, Sina Bagheri, et al.
Published: (2024)
Exploring the Maze of Multilingual Modeling
by: Nezhad, Sina Bagheri, et al.
Published: (2023)
by: Nezhad, Sina Bagheri, et al.
Published: (2023)
Enhancing Large Language Models with Neurosymbolic Reasoning for Multilingual Tasks
by: Nezhad, Sina Bagheri, et al.
Published: (2025)
by: Nezhad, Sina Bagheri, et al.
Published: (2025)
VerAs: Verify then Assess STEM Lab Reports
by: Atil, Berk, et al.
Published: (2024)
by: Atil, Berk, et al.
Published: (2024)
franciellevargas/CER: Initial Release
by: Francielle Vargas
Published: (2025)
by: Francielle Vargas
Published: (2025)
Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
by: Artemova, Ekaterina, et al.
Published: (2025)
by: Artemova, Ekaterina, et al.
Published: (2025)
SynHate: Detecting Hate Speech in Synthetic Deepfake Audio
by: Ranjan, Rishabh, et al.
Published: (2025)
by: Ranjan, Rishabh, et al.
Published: (2025)
Cross-Lingual Activation Steering for Multilingual Language Models
by: Pokharel, Rhitabrat, et al.
Published: (2026)
by: Pokharel, Rhitabrat, et al.
Published: (2026)
ThreadSumm: Summarization of Nested Discourse Threads Using Tree of Thoughts
by: Olabisi, Olubusayo, et al.
Published: (2026)
by: Olabisi, Olubusayo, et al.
Published: (2026)
Making a Long Story Short in Conversation Modeling
by: Tao, Yufei, et al.
Published: (2024)
by: Tao, Yufei, et al.
Published: (2024)
Similar Items
-
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation
by: Trager, Jackson, et al.
Published: (2025) -
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
by: Eilertsen, Brage, et al.
Published: (2025) -
Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
by: Ghorbanpour, Faeze, et al.
Published: (2025) -
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
by: Ghorbanpour, Faeze, et al.
Published: (2025) -
Explainable Semantic Textual Similarity via Dissimilar Span Detection
by: Lozano, Diego Miguel, et al.
Published: (2026)