Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster
Fuente:
arXiv
Saved in:
| Main Authors: | Calabrese, Agostina, Neves, Leonardo, Shah, Neil, Bos, Maarten W., Ross, Björn, Lapata, Mirella, Barbieri, Francesco |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compositional Generalisation for Explainable Hate Speech Detection
by: Calabrese, Agostina, et al.
Published: (2025)
by: Calabrese, Agostina, et al.
Published: (2025)
The Virality of Hate Speech on Social Media
by: Maarouf, Abdurahman, et al.
Published: (2022)
by: Maarouf, Abdurahman, et al.
Published: (2022)
K*-Means: A Parameter-free Clustering Algorithm
by: Mahon, Louis, et al.
Published: (2025)
by: Mahon, Louis, et al.
Published: (2025)
Learning to Reason for Long-Form Story Generation
by: Gurung, Alexander, et al.
Published: (2025)
by: Gurung, Alexander, et al.
Published: (2025)
Multimodal Latent Reasoning via Predictive Embeddings
by: Adhikari, Ashutosh, et al.
Published: (2026)
by: Adhikari, Ashutosh, et al.
Published: (2026)
Reasoning about Intent for Ambiguous Requests
by: Saparina, Irina, et al.
Published: (2025)
by: Saparina, Irina, et al.
Published: (2025)
A Modular Approach for Multimodal Summarization of TV Shows
by: Mahon, Louis, et al.
Published: (2024)
by: Mahon, Louis, et al.
Published: (2024)
ScreenWriter: Automatic Screenplay Generation and Movie Summarisation
by: Mahon, Louis, et al.
Published: (2024)
by: Mahon, Louis, et al.
Published: (2024)
Debating for Better Reasoning: An Unsupervised Multimodal Approach
by: Adhikari, Ashutosh, et al.
Published: (2025)
by: Adhikari, Ashutosh, et al.
Published: (2025)
AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries
by: Saparina, Irina, et al.
Published: (2024)
by: Saparina, Irina, et al.
Published: (2024)
CHIRON: Rich Character Representations in Long-Form Narratives
by: Gurung, Alexander, et al.
Published: (2024)
by: Gurung, Alexander, et al.
Published: (2024)
Integrating Large Language Models with Graph-based Reasoning for Conversational Question Answering
by: Jain, Parag, et al.
Published: (2024)
by: Jain, Parag, et al.
Published: (2024)
Disambiguate First, Parse Later: Generating Interpretations for Ambiguity Resolution in Semantic Parsing
by: Saparina, Irina, et al.
Published: (2025)
by: Saparina, Irina, et al.
Published: (2025)
Parameter-free Video Segmentation for Vision and Language Understanding
by: Mahon, Louis, et al.
Published: (2025)
by: Mahon, Louis, et al.
Published: (2025)
Context-Aware Hierarchical Merging for Long Document Summarization
by: Ou, Litu, et al.
Published: (2025)
by: Ou, Litu, et al.
Published: (2025)
Improving Generalization in Semantic Parsing by Increasing Natural Language Variation
by: Saparina, Irina, et al.
Published: (2024)
by: Saparina, Irina, et al.
Published: (2024)
Explainable AI for Hate Speech Moderation: A Stakeholder‐Centered and Sociotechnical Review
by: Muhammad Deedahwar Mazhar Qureshi, et al.
Published: (2026)
by: Muhammad Deedahwar Mazhar Qureshi, et al.
Published: (2026)
HateModerate: Testing Hate Speech Detectors against Content Moderation Policies
by: Zheng, Jiangrui, et al.
Published: (2023)
by: Zheng, Jiangrui, et al.
Published: (2023)
USE: Dynamic User Modeling with Stateful Sequence Models
by: Zhou, Zhihan, et al.
Published: (2024)
by: Zhou, Zhihan, et al.
Published: (2024)
Uncertainty Quantification in Retrieval Augmented Question Answering
by: Perez-Beltrachini, Laura, et al.
Published: (2025)
by: Perez-Beltrachini, Laura, et al.
Published: (2025)
Context-Aware Prediction of User Engagement on Online Social Platforms
by: Peters, Heinrich, et al.
Published: (2023)
by: Peters, Heinrich, et al.
Published: (2023)
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
The Enforcement and Feasibility of Hate Speech Moderation on Twitter
by: Tonneau, Manuel, et al.
Published: (2026)
by: Tonneau, Manuel, et al.
Published: (2026)
"Ignorance is Not Bliss": Designing Personalized Moderation to Address Ableist Hate on Social Media
by: Heung, Sharon, et al.
Published: (2025)
by: Heung, Sharon, et al.
Published: (2025)
Analyzing User Characteristics of Hate Speech Spreaders on Social Media
by: Geissler, Dominique, et al.
Published: (2023)
by: Geissler, Dominique, et al.
Published: (2023)
Dialogues of Dissent: Thematic and Rhetorical Dimensions of Hate and Counter-Hate Speech in Social Media Conversations
by: Levi, Effi, et al.
Published: (2025)
by: Levi, Effi, et al.
Published: (2025)
PixT3: Pixel-based Table-To-Text Generation
by: Alonso, Iñigo, et al.
Published: (2023)
by: Alonso, Iñigo, et al.
Published: (2023)
Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books
by: Papoudakis, Argyrios, et al.
Published: (2026)
by: Papoudakis, Argyrios, et al.
Published: (2026)
Finding the Right Moment: Human-Assisted Trailer Creation via Task Composition
by: Papalampidi, Pinelopi, et al.
Published: (2021)
by: Papalampidi, Pinelopi, et al.
Published: (2021)
Generating Visual Stories with Grounded and Coreferent Characters
by: Liu, Danyang, et al.
Published: (2024)
by: Liu, Danyang, et al.
Published: (2024)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
by: Gupta, Akash, et al.
Published: (2025)
by: Gupta, Akash, et al.
Published: (2025)
Lightweight Latent Reasoning for Narrative Tasks
by: Gurung, Alexander, et al.
Published: (2025)
by: Gurung, Alexander, et al.
Published: (2025)
SimLM: Can Language Models Infer Parameters of Physical Systems?
by: Memery, Sean, et al.
Published: (2023)
by: Memery, Sean, et al.
Published: (2023)
BookWorm: A Dataset for Character Description and Analysis
by: Papoudakis, Argyrios, et al.
Published: (2024)
by: Papoudakis, Argyrios, et al.
Published: (2024)
Hierarchical Indexing for Retrieval-Augmented Opinion Summarization
by: Hosking, Tom, et al.
Published: (2024)
by: Hosking, Tom, et al.
Published: (2024)
General-Purpose User Modeling with Behavioral Logs: A Snapchat Case Study
by: Fang, Qixiang, et al.
Published: (2023)
by: Fang, Qixiang, et al.
Published: (2023)
An Investigation Into Explainable Audio Hate Speech Detection
by: An, Jinmyeong, et al.
Published: (2024)
by: An, Jinmyeong, et al.
Published: (2024)
Social Media and Hate
by: Banaji, Shakuntala, et al.
Published: (2025)
by: Banaji, Shakuntala, et al.
Published: (2025)
Incorporating Human Explanations for Robust Hate Speech Detection
by: Chen, Jennifer L., et al.
Published: (2024)
by: Chen, Jennifer L., et al.
Published: (2024)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
by: Hu, Yujia, et al.
Published: (2026)
by: Hu, Yujia, et al.
Published: (2026)
Similar Items
-
Compositional Generalisation for Explainable Hate Speech Detection
by: Calabrese, Agostina, et al.
Published: (2025) -
The Virality of Hate Speech on Social Media
by: Maarouf, Abdurahman, et al.
Published: (2022) -
K*-Means: A Parameter-free Clustering Algorithm
by: Mahon, Louis, et al.
Published: (2025) -
Learning to Reason for Long-Form Story Generation
by: Gurung, Alexander, et al.
Published: (2025) -
Multimodal Latent Reasoning via Predictive Embeddings
by: Adhikari, Ashutosh, et al.
Published: (2026)