Analyzing Bias in False Refusal Behavior of Large Language Models for Hate Speech Detoxification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Im, Kyuri, Yuan, Shuzhou, Färber, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
GraSAME: Injecting Token-Level Structural Information to Pretrained Language Models via Graph-guided Self-Attention Mechanism
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
From Monolingual to Bilingual: Investigating Language Conditioning in Large Language Models for Psycholinguistic Tasks
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
Why Lift so Heavy? Slimming Large Language Models by Cutting Off the Layers
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
SimplifyMyText: An LLM-Based System for Inclusive Plain Language Text Simplification
von: Färber, Michael, et al.
Veröffentlicht: (2025)
von: Färber, Michael, et al.
Veröffentlicht: (2025)
GNNavi: Navigating the Information Flow in Large Language Models by Graph Neural Network
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
CoDAE: Adapting Large Language Models for Education via Chain-of-Thought Data Augmentation
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
PoeTone: A Framework for Constrained Generation of Structured Chinese Songci with LLMs
von: Qu, Zhan, et al.
Veröffentlicht: (2025)
von: Qu, Zhan, et al.
Veröffentlicht: (2025)
Decomposed Prompting: Probing Multilingual Linguistic Structure Knowledge in Large Language Models
von: Nie, Ercong, et al.
Veröffentlicht: (2024)
von: Nie, Ercong, et al.
Veröffentlicht: (2024)
Characterizing Selective Refusal Bias in Large Language Models
von: Khorramrouz, Adel, et al.
Veröffentlicht: (2025)
von: Khorramrouz, Adel, et al.
Veröffentlicht: (2025)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
von: Das, Amit, et al.
Veröffentlicht: (2024)
von: Das, Amit, et al.
Veröffentlicht: (2024)
The Hidden Bias: A Study on Explicit and Implicit Political Stereotypes in Large Language Models
von: Löhr, Konrad, et al.
Veröffentlicht: (2025)
von: Löhr, Konrad, et al.
Veröffentlicht: (2025)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
von: An, Bang, et al.
Veröffentlicht: (2024)
von: An, Bang, et al.
Veröffentlicht: (2024)
Can Hallucinations Help? Boosting LLMs for Drug Discovery
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
Identifying False Content and Hate Speech in Sinhala YouTube Videos by Analyzing the Audio
von: Wickramaarachchi, W. A. K. M., et al.
Veröffentlicht: (2024)
von: Wickramaarachchi, W. A. K. M., et al.
Veröffentlicht: (2024)
Tracing Relational Knowledge Recall in Large Language Models
von: Popovič, Nicholas, et al.
Veröffentlicht: (2026)
von: Popovič, Nicholas, et al.
Veröffentlicht: (2026)
Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
ToxSyn: Reducing Bias in Hate Speech Detection via Synthetic Minority Data in Brazilian Portuguese
von: Brito, Iago Alves, et al.
Veröffentlicht: (2025)
von: Brito, Iago Alves, et al.
Veröffentlicht: (2025)
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
Outcome-Constrained Large Language Models for Countering Hate Speech
von: Hong, Lingzi, et al.
Veröffentlicht: (2024)
von: Hong, Lingzi, et al.
Veröffentlicht: (2024)
HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
von: Sen, Tanmay, et al.
Veröffentlicht: (2024)
von: Sen, Tanmay, et al.
Veröffentlicht: (2024)
From Languages to Geographies: Towards Evaluating Cultural Bias in Hate Speech Datasets
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
von: Plaza-del-Arco, Flor Miriam, et al.
Veröffentlicht: (2025)
von: Plaza-del-Arco, Flor Miriam, et al.
Veröffentlicht: (2025)
Refusal Behavior in Large Language Models: A Nonlinear Perspective
von: Hildebrandt, Fabian, et al.
Veröffentlicht: (2025)
von: Hildebrandt, Fabian, et al.
Veröffentlicht: (2025)
Exploring Large Language Models for Hate Speech Detection in Rioplatense Spanish
von: Pérez, Juan Manuel, et al.
Veröffentlicht: (2024)
von: Pérez, Juan Manuel, et al.
Veröffentlicht: (2024)
Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language Models
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
DSCD: Large Language Model Detoxification with Self-Constrained Decoding
von: Dong, Ming, et al.
Veröffentlicht: (2025)
von: Dong, Ming, et al.
Veröffentlicht: (2025)
Detecting Anti-Semitic Hate Speech using Transformer-based Large Language Models
von: Liu, Dengyi, et al.
Veröffentlicht: (2024)
von: Liu, Dengyi, et al.
Veröffentlicht: (2024)
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models
von: Bui, Minh Duc, et al.
Veröffentlicht: (2024)
von: Bui, Minh Duc, et al.
Veröffentlicht: (2024)
Graph-Guided Textual Explanation Generation Framework
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks
von: Ma, Bolei, et al.
Veröffentlicht: (2024)
von: Ma, Bolei, et al.
Veröffentlicht: (2024)
Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models
von: Schreieder, Tobias, et al.
Veröffentlicht: (2025)
von: Schreieder, Tobias, et al.
Veröffentlicht: (2025)
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech
von: Ludwig, Florian, et al.
Veröffentlicht: (2025)
von: Ludwig, Florian, et al.
Veröffentlicht: (2025)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2024)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2024)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
von: Jain, Neel, et al.
Veröffentlicht: (2024)
von: Jain, Neel, et al.
Veröffentlicht: (2024)
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages
von: Muhammad, Shamsuddeen Hassan, et al.
Veröffentlicht: (2025)
von: Muhammad, Shamsuddeen Hassan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025) -
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025) -
GraSAME: Injecting Token-Level Structural Information to Pretrained Language Models via Graph-guided Self-Attention Mechanism
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024) -
From Monolingual to Bilingual: Investigating Language Conditioning in Large Language Models for Psycholinguistic Tasks
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025) -
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)