Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yifan, Jobanputra, Mayank, Lee, Ji-Ung, Oh, Soyoung, Valera, Isabel, Demberg, Vera |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
B-cos LM: Efficiently Transforming Pre-trained Language Models for Improved Explainability
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025)
by: Bauer, Julie, et al.
Published: (2025)
Tug-of-war between idioms' figurative and literal interpretations in LLMs
by: Oh, Soyoung, et al.
Published: (2025)
by: Oh, Soyoung, et al.
Published: (2025)
RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation Framework
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
An Investigation Into Explainable Audio Hate Speech Detection
by: An, Jinmyeong, et al.
Published: (2024)
by: An, Jinmyeong, et al.
Published: (2024)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
by: Hu, Yujia, et al.
Published: (2026)
by: Hu, Yujia, et al.
Published: (2026)
An Effective, Robust and Fairness-aware Hate Speech Detection Framework
by: Mou, Guanyi, et al.
Published: (2024)
by: Mou, Guanyi, et al.
Published: (2024)
Hate Speech Detection with Generalizable Target-aware Fairness
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
Incorporating Distributions of Discourse Structure for Long Document Abstractive Summarization
by: Liu, Dongqi, et al.
Published: (2023)
by: Liu, Dongqi, et al.
Published: (2023)
Can LLMs subtract numbers?
by: Jobanputra, Mayank, et al.
Published: (2025)
by: Jobanputra, Mayank, et al.
Published: (2025)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
by: Jobanputra, Mayank, et al.
Published: (2025)
by: Jobanputra, Mayank, et al.
Published: (2025)
Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection
by: Kim, Yumin, et al.
Published: (2025)
by: Kim, Yumin, et al.
Published: (2025)
xList-Hate: A Checklist-Based Framework for Interpretable and Generalizable Hate Speech Detection
by: Girón, Adrián, et al.
Published: (2026)
by: Girón, Adrián, et al.
Published: (2026)
ChatGPT vs Human-authored Text: Insights into Controllable Text Summarization and Sentence Style Transfer
by: Liu, Dongqi, et al.
Published: (2023)
by: Liu, Dongqi, et al.
Published: (2023)
RST-LoRA: A Discourse-Aware Low-Rank Adaptation for Long Document Abstractive Summarization
by: Liu, Dongqi, et al.
Published: (2024)
by: Liu, Dongqi, et al.
Published: (2024)
SciNews: From Scholarly Complexities to Public Narratives -- A Dataset for Scientific News Report Generation
by: Liu, Dongqi, et al.
Published: (2024)
by: Liu, Dongqi, et al.
Published: (2024)
Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI
by: Böck, Adrian Jaques, et al.
Published: (2024)
by: Böck, Adrian Jaques, et al.
Published: (2024)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
by: Proskurina, Irina, et al.
Published: (2025)
by: Proskurina, Irina, et al.
Published: (2025)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Understanding Fairness-Accuracy Trade-offs in Machine Learning Models: Does Promoting Fairness Undermine Performance?
by: Liu, Junhua, et al.
Published: (2024)
by: Liu, Junhua, et al.
Published: (2024)
Compositional Generalisation for Explainable Hate Speech Detection
by: Calabrese, Agostina, et al.
Published: (2025)
by: Calabrese, Agostina, et al.
Published: (2025)
Human Speech Perception in Noise: Can Large Language Models Paraphrase to Improve It?
by: Chingacham, Anupama, et al.
Published: (2024)
by: Chingacham, Anupama, et al.
Published: (2024)
Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"?
by: Liu, Tong, et al.
Published: (2023)
by: Liu, Tong, et al.
Published: (2023)
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
by: Nghiem, Huy, et al.
Published: (2024)
by: Nghiem, Huy, et al.
Published: (2024)
Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
by: Gajewska, Ewelina, et al.
Published: (2025)
by: Gajewska, Ewelina, et al.
Published: (2025)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
by: Park, Junyeong, et al.
Published: (2025)
by: Park, Junyeong, et al.
Published: (2025)
Incorporating Human Explanations for Robust Hate Speech Detection
by: Chen, Jennifer L., et al.
Published: (2024)
by: Chen, Jennifer L., et al.
Published: (2024)
Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster
by: Calabrese, Agostina, et al.
Published: (2024)
by: Calabrese, Agostina, et al.
Published: (2024)
Explainable Speech Emotion Recognition: Weighted Attribute Fairness to Model Demographic Contributions to Social Bias
by: Ogunnubi, Tomisin, et al.
Published: (2026)
by: Ogunnubi, Tomisin, et al.
Published: (2026)
Fair Summarization: Bridging Quality and Diversity in Extractive Summaries
by: Nezhad, Sina Bagheri, et al.
Published: (2024)
by: Nezhad, Sina Bagheri, et al.
Published: (2024)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
by: Lee, Nayeon, et al.
Published: (2023)
by: Lee, Nayeon, et al.
Published: (2023)
Analysis and Detection of Multilingual Hate Speech Using Transformer Based Deep Learning
by: Das, Arijit, et al.
Published: (2024)
by: Das, Arijit, et al.
Published: (2024)
SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia
by: Ng, Ri Chi, et al.
Published: (2026)
by: Ng, Ri Chi, et al.
Published: (2026)
Deep Knowledge-Infusion For Explainable Depression Detection
by: Dalal, Sumit, et al.
Published: (2024)
by: Dalal, Sumit, et al.
Published: (2024)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
by: Kumarage, Tharindu, et al.
Published: (2024)
by: Kumarage, Tharindu, et al.
Published: (2024)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
by: Almohaimeed, Saad, et al.
Published: (2025)
by: Almohaimeed, Saad, et al.
Published: (2025)
Prompting Implicit Discourse Relation Annotation
by: Yung, Frances, et al.
Published: (2024)
by: Yung, Frances, et al.
Published: (2024)
Fairness of Automatic Speech Recognition: Looking Through a Philosophical Lens
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech
by: Ludwig, Florian, et al.
Published: (2025)
by: Ludwig, Florian, et al.
Published: (2025)
Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages
by: Prome, Ruhina Tabasshum, et al.
Published: (2025)
by: Prome, Ruhina Tabasshum, et al.
Published: (2025)
Similar Items
-
B-cos LM: Efficiently Transforming Pre-trained Language Models for Improved Explainability
by: Wang, Yifan, et al.
Published: (2025) -
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025) -
Tug-of-war between idioms' figurative and literal interpretations in LLMs
by: Oh, Soyoung, et al.
Published: (2025) -
RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation Framework
by: Wang, Yifan, et al.
Published: (2024) -
An Investigation Into Explainable Audio Hate Speech Detection
by: An, Jinmyeong, et al.
Published: (2024)