Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Piot, Paloma, Martín-Rodilla, Patricia, Parapar, Javier |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2024)
by: Piot, Paloma, et al.
Published: (2024)
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
by: Piot, Paloma, et al.
Published: (2024)
by: Piot, Paloma, et al.
Published: (2024)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Towards Efficient and Explainable Hate Speech Detection via Model Distillation
by: Piot, Paloma, et al.
Published: (2024)
by: Piot, Paloma, et al.
Published: (2024)
Bridging Gaps in Hate Speech Detection: Meta-Collections and Benchmarks for Low-Resource Iberian Languages
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
WATCHED: A Web AI Agent Tool for Combating Hate Speech by Expanding Data
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
HateDebias: On the Diversity and Variability of Hate Speech Debiasing
by: Wu, Hongyan, et al.
Published: (2024)
by: Wu, Hongyan, et al.
Published: (2024)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
by: Das, Amit, et al.
Published: (2024)
by: Das, Amit, et al.
Published: (2024)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
by: Zahid, Anwar Hossain, et al.
Published: (2025)
by: Zahid, Anwar Hossain, et al.
Published: (2025)
Analyzing Bias in False Refusal Behavior of Large Language Models for Hate Speech Detoxification
by: Im, Kyuri, et al.
Published: (2026)
by: Im, Kyuri, et al.
Published: (2026)
Detecting Anti-Semitic Hate Speech using Transformer-based Large Language Models
by: Liu, Dengyi, et al.
Published: (2024)
by: Liu, Dengyi, et al.
Published: (2024)
HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
by: Sen, Tanmay, et al.
Published: (2024)
by: Sen, Tanmay, et al.
Published: (2024)
Exploring Large Language Models for Hate Speech Detection in Rioplatense Spanish
by: Pérez, Juan Manuel, et al.
Published: (2024)
by: Pérez, Juan Manuel, et al.
Published: (2024)
Hate Speech Detection using Large Language Models with Data Augmentation and Feature Enhancement
by: Nge, Brian Jing Hong, et al.
Published: (2026)
by: Nge, Brian Jing Hong, et al.
Published: (2026)
Self-Debias: Self-correcting for Debiasing Large Language Models
by: Feng, Xuan, et al.
Published: (2026)
by: Feng, Xuan, et al.
Published: (2026)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
by: Nirmal, Ayushi, et al.
Published: (2024)
by: Nirmal, Ayushi, et al.
Published: (2024)
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models
by: Bui, Minh Duc, et al.
Published: (2024)
by: Bui, Minh Duc, et al.
Published: (2024)
Outcome-Constrained Large Language Models for Countering Hate Speech
by: Hong, Lingzi, et al.
Published: (2024)
by: Hong, Lingzi, et al.
Published: (2024)
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech
by: Ludwig, Florian, et al.
Published: (2025)
by: Ludwig, Florian, et al.
Published: (2025)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
by: Kumarage, Tharindu, et al.
Published: (2024)
by: Kumarage, Tharindu, et al.
Published: (2024)
From Languages to Geographies: Towards Evaluating Cultural Bias in Hate Speech Datasets
by: Tonneau, Manuel, et al.
Published: (2024)
by: Tonneau, Manuel, et al.
Published: (2024)
An Investigation of Large Language Models for Real-World Hate Speech Detection
by: Guo, Keyan, et al.
Published: (2024)
by: Guo, Keyan, et al.
Published: (2024)
Explainable Depression Symptom Detection in Social Media
by: Bao, Eliseo, et al.
Published: (2023)
by: Bao, Eliseo, et al.
Published: (2023)
From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
by: Cai, Zefan, et al.
Published: (2025)
by: Cai, Zefan, et al.
Published: (2025)
Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs
by: Sidibomma, Rushendra, et al.
Published: (2024)
by: Sidibomma, Rushendra, et al.
Published: (2024)
ReDSM5: A Reddit Dataset for DSM-5 Depression Detection
by: Bao, Eliseo, et al.
Published: (2025)
by: Bao, Eliseo, et al.
Published: (2025)
Prompting Fairness: Integrating Causality to Debias Large Language Models
by: Li, Jingling, et al.
Published: (2024)
by: Li, Jingling, et al.
Published: (2024)
DebiasRAG: A Tuning-Free Path to Fair Generation in Large Language Models through Retrieval-Augmented Generation
by: Chu, Rui, et al.
Published: (2026)
by: Chu, Rui, et al.
Published: (2026)
"Is Hate Lost in Translation?": Evaluation of Multilingual LGBTQIA+ Hate Speech Detection
by: Chan, Fai Leui, et al.
Published: (2024)
by: Chan, Fai Leui, et al.
Published: (2024)
EkoHate: Abusive Language and Hate Speech Detection for Code-switched Political Discussions on Nigerian Twitter
by: Ilevbare, Comfort Eseohen, et al.
Published: (2024)
by: Ilevbare, Comfort Eseohen, et al.
Published: (2024)
MasonPerplexity at Multimodal Hate Speech Event Detection 2024: Hate Speech and Target Detection Using Transformer Ensembles
by: Ganguly, Amrita, et al.
Published: (2024)
by: Ganguly, Amrita, et al.
Published: (2024)
Efficient Hate Speech Detection: A Three-Layer LoRA-Tuned BERTweet Framework
by: El-Bahnasawi, Mahmoud
Published: (2025)
by: El-Bahnasawi, Mahmoud
Published: (2025)
xList-Hate: A Checklist-Based Framework for Interpretable and Generalizable Hate Speech Detection
by: Girón, Adrián, et al.
Published: (2026)
by: Girón, Adrián, et al.
Published: (2026)
PartisanLens: A Multilingual Dataset of Hyperpartisan and Conspiratorial Immigration Narratives in European Media
by: Maggini, Michele Joshua, et al.
Published: (2026)
by: Maggini, Michele Joshua, et al.
Published: (2026)
Beyond the Explicit: A Bilingual Dataset for Dehumanization Detection in Social Media
by: Assenmacher, Dennis, et al.
Published: (2025)
by: Assenmacher, Dennis, et al.
Published: (2025)
Reducing Large Language Model Bias with Emphasis on 'Restricted Industries': Automated Dataset Augmentation and Prejudice Quantification
by: Mondal, Devam, et al.
Published: (2024)
by: Mondal, Devam, et al.
Published: (2024)
Compositional Generalisation for Explainable Hate Speech Detection
by: Calabrese, Agostina, et al.
Published: (2025)
by: Calabrese, Agostina, et al.
Published: (2025)
Automatic Textual Normalization for Hate Speech Detection
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2023)
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2023)
Similar Items
-
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2024) -
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
by: Piot, Paloma, et al.
Published: (2024) -
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2025) -
Towards Efficient and Explainable Hate Speech Detection via Model Distillation
by: Piot, Paloma, et al.
Published: (2024) -
Bridging Gaps in Hate Speech Detection: Meta-Collections and Benchmarks for Low-Resource Iberian Languages
by: Piot, Paloma, et al.
Published: (2025)