Hate Personified: Investigating the role of LLMs in content moderation
Fuente:
arXiv
Saved in:
| Main Authors: | Masud, Sarah, Singh, Sahajpreet, Hangya, Viktor, Fraser, Alexander, Chakraborty, Tanmoy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Focal Inferential Infusion Coupled with Tractable Density Discrimination for Implicit Hate Detection
by: Masud, Sarah, et al.
Published: (2023)
by: Masud, Sarah, et al.
Published: (2023)
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have
by: Hangya, Viktor, et al.
Published: (2023)
by: Hangya, Viktor, et al.
Published: (2023)
Style-Specific Neurons for Steering LLMs in Text Style Transfer
by: Lai, Wen, et al.
Published: (2024)
by: Lai, Wen, et al.
Published: (2024)
SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes
by: Nandi, Palash, et al.
Published: (2024)
by: Nandi, Palash, et al.
Published: (2024)
Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
by: Singh, Sahajpreet, et al.
Published: (2025)
by: Singh, Sahajpreet, et al.
Published: (2025)
Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
by: Singh, Daman Deep, et al.
Published: (2025)
by: Singh, Daman Deep, et al.
Published: (2025)
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
GitSearch: Enhancing Community Notes Generation with Gap-Informed Targeted Search
by: Singh, Sahajpreet, et al.
Published: (2026)
by: Singh, Sahajpreet, et al.
Published: (2026)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
by: Jin, Yiping, et al.
Published: (2024)
by: Jin, Yiping, et al.
Published: (2024)
Probing Critical Learning Dynamics of PLMs for Hate Speech Detection
by: Masud, Sarah, et al.
Published: (2024)
by: Masud, Sarah, et al.
Published: (2024)
Extending Multilingual Machine Translation through Imitation Learning
by: Lai, Wen, et al.
Published: (2023)
by: Lai, Wen, et al.
Published: (2023)
Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models
by: Sharma, Shivam, et al.
Published: (2025)
by: Sharma, Shivam, et al.
Published: (2025)
PersLLM: A Personified Training Approach for Large Language Models
by: Zeng, Zheni, et al.
Published: (2024)
by: Zeng, Zheni, et al.
Published: (2024)
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech
by: Yadav, Neemesh, et al.
Published: (2024)
by: Yadav, Neemesh, et al.
Published: (2024)
Information Anxiety in Large Language Models
by: Bajpai, Prasoon, et al.
Published: (2024)
by: Bajpai, Prasoon, et al.
Published: (2024)
CommunityFact: A Dynamic, Multilingual, Multi-domain Benchmark for Misinformation Detection in the Wild
by: Singh, Sahajpreet, et al.
Published: (2026)
by: Singh, Sahajpreet, et al.
Published: (2026)
Independent fact-checking organizations exhibit a departure from political neutrality
by: Singh, Sahajpreet, et al.
Published: (2024)
by: Singh, Sahajpreet, et al.
Published: (2024)
AI Across Borders: Exploring Perceptions and Interactions in Higher Education
by: Gerard, Juliana, et al.
Published: (2024)
by: Gerard, Juliana, et al.
Published: (2024)
SUKHSANDESH: An Avatar Therapeutic Question Answering Platform for Sexual Education in Rural India
by: Singh, Salam Michael, et al.
Published: (2024)
by: Singh, Salam Michael, et al.
Published: (2024)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
by: Agarwal, Siddhant, et al.
Published: (2024)
by: Agarwal, Siddhant, et al.
Published: (2024)
Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
by: Gajewska, Ewelina, et al.
Published: (2025)
by: Gajewska, Ewelina, et al.
Published: (2025)
The Unseen Targets of Hate -- A Systematic Review of Hateful Communication Datasets
by: Yu, Zehui, et al.
Published: (2024)
by: Yu, Zehui, et al.
Published: (2024)
CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs
by: Hengle, Amey, et al.
Published: (2025)
by: Hengle, Amey, et al.
Published: (2025)
From Understanding to Generation: An Efficient Shortcut for Evaluating Language Models
by: Hangya, Viktor, et al.
Published: (2025)
by: Hangya, Viktor, et al.
Published: (2025)
The Enforcement and Feasibility of Hate Speech Moderation on Twitter
by: Tonneau, Manuel, et al.
Published: (2026)
by: Tonneau, Manuel, et al.
Published: (2026)
Towards Weakly-Supervised Hate Speech Classification Across Datasets
by: Jin, Yiping, et al.
Published: (2023)
by: Jin, Yiping, et al.
Published: (2023)
The Psychology of Falsehood: A Human-Centric Survey of Misinformation Detection
by: Nandi, Arghodeep, et al.
Published: (2025)
by: Nandi, Arghodeep, et al.
Published: (2025)
An Investigation of Large Language Models for Real-World Hate Speech Detection
by: Guo, Keyan, et al.
Published: (2024)
by: Guo, Keyan, et al.
Published: (2024)
Multilingualism, Transnationality, and K-pop in the Online #StopAsianHate Movement
by: Masis, Tessa, et al.
Published: (2025)
by: Masis, Tessa, et al.
Published: (2025)
Few-shot Hate Speech Detection Based on the MindSpore Framework
by: Qin, Zhenkai, et al.
Published: (2025)
by: Qin, Zhenkai, et al.
Published: (2025)
Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation
by: Herrmann, Nils A., et al.
Published: (2026)
by: Herrmann, Nils A., et al.
Published: (2026)
Measuring Online Hate on 4chan using Pre-trained Deep Learning Models
by: Bermudez-Villalva, Adrian, et al.
Published: (2025)
by: Bermudez-Villalva, Adrian, et al.
Published: (2025)
Sometimes the Model doth Preach: Quantifying Religious Bias in Open LLMs through Demographic Analysis in Asian Nations
by: Shankar, Hari, et al.
Published: (2025)
by: Shankar, Hari, et al.
Published: (2025)
A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance
by: Melis, Matteo, et al.
Published: (2025)
by: Melis, Matteo, et al.
Published: (2025)
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025)
by: Bauer, Julie, et al.
Published: (2025)
Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
Somatic in the East, Psychological in the West?: Investigating Clinically-Grounded Cross-Cultural Depression Symptom Expression in LLMs
by: Sakai, Shintaro, et al.
Published: (2025)
by: Sakai, Shintaro, et al.
Published: (2025)
Exploring LLMs for Predicting Tutor Strategy and Student Outcomes in Dialogues
by: Ikram, Fareya, et al.
Published: (2025)
by: Ikram, Fareya, et al.
Published: (2025)
Similar Items
-
Focal Inferential Infusion Coupled with Tractable Density Discrimination for Implicit Hate Detection
by: Masud, Sarah, et al.
Published: (2023) -
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have
by: Hangya, Viktor, et al.
Published: (2023) -
Style-Specific Neurons for Steering LLMs in Text Style Transfer
by: Lai, Wen, et al.
Published: (2024) -
SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes
by: Nandi, Palash, et al.
Published: (2024) -
Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
by: Singh, Sahajpreet, et al.
Published: (2025)