Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Gajewska, Ewelina, Derbent, Arda, Chudziak, Jaroslaw A, Budzynska, Katarzyna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
by: Gajewska, Ewelina, et al.
Published: (2026)
by: Gajewska, Ewelina, et al.
Published: (2026)
Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
by: Gajewska, Ewelina, et al.
Published: (2026)
by: Gajewska, Ewelina, et al.
Published: (2026)
Leveraging a Multi-Agent LLM-Based System to Educate Teachers in Hate Incidents Management
by: Gajewska, Ewelina, et al.
Published: (2025)
by: Gajewska, Ewelina, et al.
Published: (2025)
Predicting the winner of the US 2024 elections using trust analytics
by: Budzynska, Katarzyna, et al.
Published: (2024)
by: Budzynska, Katarzyna, et al.
Published: (2024)
A Natural Language Agentic Approach to Study Affective Polarization
by: Malvicini, Stephanie Anneris, et al.
Published: (2026)
by: Malvicini, Stephanie Anneris, et al.
Published: (2026)
The Lovelace Test of Intelligence: Can Humans Recognise and Esteem AI-Generated Art?
by: Gajewska, Ewelina
Published: (2025)
by: Gajewska, Ewelina
Published: (2025)
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025)
by: Bauer, Julie, et al.
Published: (2025)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
by: Jin, Yiping, et al.
Published: (2024)
by: Jin, Yiping, et al.
Published: (2024)
Ethos and Pathos in Online Group Discussions: Corpora for Polarisation Issues in Social Media
by: Gajewska, Ewelina, et al.
Published: (2024)
by: Gajewska, Ewelina, et al.
Published: (2024)
Heterogeneous Debate Engine: Identity-Grounded Cognitive Architecture for Resilient LLM-Based Ethical Tutoring
by: Masłowski, Jakub, et al.
Published: (2026)
by: Masłowski, Jakub, et al.
Published: (2026)
Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
by: Singh, Daman Deep, et al.
Published: (2025)
by: Singh, Daman Deep, et al.
Published: (2025)
Human-Centric NLP or AI-Centric Illusion?: A Critical Investigation
by: Spencer, Piyapath T
Published: (2024)
by: Spencer, Piyapath T
Published: (2024)
Few-shot Hate Speech Detection Based on the MindSpore Framework
by: Qin, Zhenkai, et al.
Published: (2025)
by: Qin, Zhenkai, et al.
Published: (2025)
The Enforcement and Feasibility of Hate Speech Moderation on Twitter
by: Tonneau, Manuel, et al.
Published: (2026)
by: Tonneau, Manuel, et al.
Published: (2026)
SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes
by: Nandi, Palash, et al.
Published: (2024)
by: Nandi, Palash, et al.
Published: (2024)
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation
by: Herrmann, Nils A., et al.
Published: (2026)
by: Herrmann, Nils A., et al.
Published: (2026)
Diagnosing Hate Speech Classification: Where Do Humans and Machines Disagree, and Why?
by: Yang, Xilin
Published: (2024)
by: Yang, Xilin
Published: (2024)
Hate Personified: Investigating the role of LLMs in content moderation
by: Masud, Sarah, et al.
Published: (2024)
by: Masud, Sarah, et al.
Published: (2024)
Deep Learning Approaches for Detecting Adversarial Cyberbullying and Hate Speech in Social Networks
by: Azumah, Sylvia Worlali, et al.
Published: (2024)
by: Azumah, Sylvia Worlali, et al.
Published: (2024)
Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
The Psychology of Falsehood: A Human-Centric Survey of Misinformation Detection
by: Nandi, Arghodeep, et al.
Published: (2025)
by: Nandi, Arghodeep, et al.
Published: (2025)
Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
Towards Weakly-Supervised Hate Speech Classification Across Datasets
by: Jin, Yiping, et al.
Published: (2023)
by: Jin, Yiping, et al.
Published: (2023)
Addressing Both Statistical and Causal Gender Fairness in NLP Models
by: Chen, Hannah, et al.
Published: (2024)
by: Chen, Hannah, et al.
Published: (2024)
A Comprehensive Study on NLP Data Augmentation for Hate Speech Detection: Legacy Methods, BERT, and LLMs
by: Jahan, Md Saroar, et al.
Published: (2024)
by: Jahan, Md Saroar, et al.
Published: (2024)
The Unseen Targets of Hate -- A Systematic Review of Hateful Communication Datasets
by: Yu, Zehui, et al.
Published: (2024)
by: Yu, Zehui, et al.
Published: (2024)
A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance
by: Melis, Matteo, et al.
Published: (2025)
by: Melis, Matteo, et al.
Published: (2025)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
by: Ngueajio, Mikel K., et al.
Published: (2025)
by: Ngueajio, Mikel K., et al.
Published: (2025)
Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
by: Xian, Ruicheng, et al.
Published: (2025)
by: Xian, Ruicheng, et al.
Published: (2025)
TACLA: An LLM-Based Multi-Agent Tool for Transactional Analysis Training in Education
by: Zamojska, Monika, et al.
Published: (2025)
by: Zamojska, Monika, et al.
Published: (2025)
Focal Inferential Infusion Coupled with Tractable Density Discrimination for Implicit Hate Detection
by: Masud, Sarah, et al.
Published: (2023)
by: Masud, Sarah, et al.
Published: (2023)
Generalizing Hate Speech Detection Using Multi-Task Learning: A Case Study of Political Public Figures
by: Yuan, Lanqin, et al.
Published: (2022)
by: Yuan, Lanqin, et al.
Published: (2022)
Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
by: Prama, Tabia Tanzin, et al.
Published: (2025)
by: Prama, Tabia Tanzin, et al.
Published: (2025)
GAIus: Combining Genai with Legal Clauses Retrieval for Knowledge-based Assistant
by: Matak, Michał, et al.
Published: (2025)
by: Matak, Michał, et al.
Published: (2025)
Explainable Rule Application via Structured Prompting: A Neural-Symbolic Approach
by: Sadowski, Albert, et al.
Published: (2025)
by: Sadowski, Albert, et al.
Published: (2025)
On Verifiable Legal Reasoning: A Multi-Agent Framework with Formalized Knowledge Representations
by: Sadowski, Albert, et al.
Published: (2025)
by: Sadowski, Albert, et al.
Published: (2025)
Multi-Agent Dialectical Refinement for Enhanced Argument Classification
by: Bąba, Jakub, et al.
Published: (2026)
by: Bąba, Jakub, et al.
Published: (2026)
An Investigation of Large Language Models for Real-World Hate Speech Detection
by: Guo, Keyan, et al.
Published: (2024)
by: Guo, Keyan, et al.
Published: (2024)
Similar Items
-
Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
by: Gajewska, Ewelina, et al.
Published: (2026) -
Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
by: Gajewska, Ewelina, et al.
Published: (2026) -
Leveraging a Multi-Agent LLM-Based System to Educate Teachers in Hate Incidents Management
by: Gajewska, Ewelina, et al.
Published: (2025) -
Predicting the winner of the US 2024 elections using trust analytics
by: Budzynska, Katarzyna, et al.
Published: (2024) -
A Natural Language Agentic Approach to Study Affective Polarization
by: Malvicini, Stephanie Anneris, et al.
Published: (2026)