Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Shuzhou, Nie, Ercong, Tawfelis, Mario, Schmid, Helmut, Schütze, Hinrich, Färber, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
GNNavi: Navigating the Information Flow in Large Language Models by Graph Neural Network
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
Analyzing Bias in False Refusal Behavior of Large Language Models for Hate Speech Detoxification
von: Im, Kyuri, et al.
Veröffentlicht: (2026)
von: Im, Kyuri, et al.
Veröffentlicht: (2026)
Decomposed Prompting: Probing Multilingual Linguistic Structure Knowledge in Large Language Models
von: Nie, Ercong, et al.
Veröffentlicht: (2024)
von: Nie, Ercong, et al.
Veröffentlicht: (2024)
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
von: Nie, Ercong, et al.
Veröffentlicht: (2025)
von: Nie, Ercong, et al.
Veröffentlicht: (2025)
From Monolingual to Bilingual: Investigating Language Conditioning in Large Language Models for Psycholinguistic Tasks
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks
von: Ma, Bolei, et al.
Veröffentlicht: (2024)
von: Ma, Bolei, et al.
Veröffentlicht: (2024)
Why Lift so Heavy? Slimming Large Language Models by Cutting Off the Layers
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024)
HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
von: Sen, Tanmay, et al.
Veröffentlicht: (2024)
von: Sen, Tanmay, et al.
Veröffentlicht: (2024)
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
A Federated Approach to Few-Shot Hate Speech Detection for Marginalized Communities
von: Ye, Haotian, et al.
Veröffentlicht: (2024)
von: Ye, Haotian, et al.
Veröffentlicht: (2024)
Large Language Models as Neurolinguistic Subjects: Discrepancy between Performance and Competence
von: He, Linyang, et al.
Veröffentlicht: (2024)
von: He, Linyang, et al.
Veröffentlicht: (2024)
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models
von: Bui, Minh Duc, et al.
Veröffentlicht: (2024)
von: Bui, Minh Duc, et al.
Veröffentlicht: (2024)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
von: Das, Amit, et al.
Veröffentlicht: (2024)
von: Das, Amit, et al.
Veröffentlicht: (2024)
An Investigation of Large Language Models for Real-World Hate Speech Detection
von: Guo, Keyan, et al.
Veröffentlicht: (2024)
von: Guo, Keyan, et al.
Veröffentlicht: (2024)
CoDAE: Adapting Large Language Models for Education via Chain-of-Thought Data Augmentation
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
"Is Hate Lost in Translation?": Evaluation of Multilingual LGBTQIA+ Hate Speech Detection
von: Chan, Fai Leui, et al.
Veröffentlicht: (2024)
von: Chan, Fai Leui, et al.
Veröffentlicht: (2024)
HateDebias: On the Diversity and Variability of Hate Speech Debiasing
von: Wu, Hongyan, et al.
Veröffentlicht: (2024)
von: Wu, Hongyan, et al.
Veröffentlicht: (2024)
Exploring Large Language Models for Hate Speech Detection in Rioplatense Spanish
von: Pérez, Juan Manuel, et al.
Veröffentlicht: (2024)
von: Pérez, Juan Manuel, et al.
Veröffentlicht: (2024)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2024)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2024)
Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
von: Chapagain, Santosh, et al.
Veröffentlicht: (2025)
von: Chapagain, Santosh, et al.
Veröffentlicht: (2025)
Seeing Hate Differently: Hate Subspace Modeling for Culture-Aware Hate Speech Detection
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
Outcome-Constrained Large Language Models for Countering Hate Speech
von: Hong, Lingzi, et al.
Veröffentlicht: (2024)
von: Hong, Lingzi, et al.
Veröffentlicht: (2024)
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
Recent Advances in Hate Speech Moderation: Multimodality and the Role of Large Models
von: Hee, Ming Shan, et al.
Veröffentlicht: (2024)
von: Hee, Ming Shan, et al.
Veröffentlicht: (2024)
Demystifying Hateful Content: Leveraging Large Multimodal Models for Hateful Meme Detection with Explainable Decisions
von: Hee, Ming Shan, et al.
Veröffentlicht: (2025)
von: Hee, Ming Shan, et al.
Veröffentlicht: (2025)
EkoHate: Abusive Language and Hate Speech Detection for Code-switched Political Discussions on Nigerian Twitter
von: Ilevbare, Comfort Eseohen, et al.
Veröffentlicht: (2024)
von: Ilevbare, Comfort Eseohen, et al.
Veröffentlicht: (2024)
When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
von: Ocampo, Nicolás Benjamín, et al.
Veröffentlicht: (2026)
NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative Data
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
von: Tonneau, Manuel, et al.
Veröffentlicht: (2024)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
von: Proskurina, Irina, et al.
Veröffentlicht: (2025)
von: Proskurina, Irina, et al.
Veröffentlicht: (2025)
ViHateT5: Enhancing Hate Speech Detection in Vietnamese With A Unified Text-to-Text Transformer Model
von: Nguyen, Luan Thanh
Veröffentlicht: (2024)
von: Nguyen, Luan Thanh
Veröffentlicht: (2024)
An Investigation Into Explainable Audio Hate Speech Detection
von: An, Jinmyeong, et al.
Veröffentlicht: (2024)
von: An, Jinmyeong, et al.
Veröffentlicht: (2024)
Web(er) of Hate: A Survey on How Hate Speech Is Typed
von: Wang, Luna, et al.
Veröffentlicht: (2025)
von: Wang, Luna, et al.
Veröffentlicht: (2025)
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages
von: Muhammad, Shamsuddeen Hassan, et al.
Veröffentlicht: (2025)
von: Muhammad, Shamsuddeen Hassan, et al.
Veröffentlicht: (2025)
BMIKE-53: Investigating Cross-Lingual Knowledge Editing with In-Context Learning
von: Nie, Ercong, et al.
Veröffentlicht: (2024)
von: Nie, Ercong, et al.
Veröffentlicht: (2024)
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech
von: Ludwig, Florian, et al.
Veröffentlicht: (2025)
von: Ludwig, Florian, et al.
Veröffentlicht: (2025)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
von: Jin, Yiping, et al.
Veröffentlicht: (2024)
von: Jin, Yiping, et al.
Veröffentlicht: (2024)
Detecting Anti-Semitic Hate Speech using Transformer-based Large Language Models
von: Liu, Dengyi, et al.
Veröffentlicht: (2024)
von: Liu, Dengyi, et al.
Veröffentlicht: (2024)
MasonPerplexity at Multimodal Hate Speech Event Detection 2024: Hate Speech and Target Detection Using Transformer Ensembles
von: Ganguly, Amrita, et al.
Veröffentlicht: (2024)
von: Ganguly, Amrita, et al.
Veröffentlicht: (2024)
Compositional Generalisation for Explainable Hate Speech Detection
von: Calabrese, Agostina, et al.
Veröffentlicht: (2025)
von: Calabrese, Agostina, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025) -
GNNavi: Navigating the Information Flow in Large Language Models by Graph Neural Network
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2024) -
Analyzing Bias in False Refusal Behavior of Large Language Models for Hate Speech Detoxification
von: Im, Kyuri, et al.
Veröffentlicht: (2026) -
Decomposed Prompting: Probing Multilingual Linguistic Structure Knowledge in Large Language Models
von: Nie, Ercong, et al.
Veröffentlicht: (2024) -
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
von: Nie, Ercong, et al.
Veröffentlicht: (2025)