Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness
Fuente:
arXiv
Saved in:
| Main Authors: | Alipour, Shayan, Sen, Indira, Samory, Mattia, Mitra, Tanushree |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Gray Area: Characterizing Moderator Disagreement on Reddit
by: Alipour, Shayan, et al.
Published: (2026)
by: Alipour, Shayan, et al.
Published: (2026)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
by: Sen, Indira, et al.
Published: (2023)
by: Sen, Indira, et al.
Published: (2023)
The Unseen Targets of Hate -- A Systematic Review of Hateful Communication Datasets
by: Yu, Zehui, et al.
Published: (2024)
by: Yu, Zehui, et al.
Published: (2024)
The Implications of Open Generative Models in Human-Centered Data Science Work: A Case Study with Fact-Checking Organizations
by: Wolfe, Robert, et al.
Published: (2024)
by: Wolfe, Robert, et al.
Published: (2024)
A Multilingual Similarity Dataset for News Article Frame
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation
by: Samory, Mattia, et al.
Published: (2025)
by: Samory, Mattia, et al.
Published: (2025)
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
From Measurement Instruments to Data: Leveraging Theory-Driven Synthetic Training Data for Classifying Social Constructs
by: Birkenmaier, Lukas, et al.
Published: (2024)
by: Birkenmaier, Lukas, et al.
Published: (2024)
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations
by: Xiao, Yunze, et al.
Published: (2024)
by: Xiao, Yunze, et al.
Published: (2024)
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
by: Shen, Hua, et al.
Published: (2025)
by: Shen, Hua, et al.
Published: (2025)
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs
by: Sen, Indira, et al.
Published: (2025)
by: Sen, Indira, et al.
Published: (2025)
Tell Me What You Know About Sexism: Expert-LLM Interaction Strategies and Co-Created Definitions for Zero-Shot Sexism Detection
by: Reuver, Myrthe, et al.
Published: (2025)
by: Reuver, Myrthe, et al.
Published: (2025)
Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language Models
by: Faulborn, Mats, et al.
Published: (2025)
by: Faulborn, Mats, et al.
Published: (2025)
ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios
by: Phutane, Mahika, et al.
Published: (2025)
by: Phutane, Mahika, et al.
Published: (2025)
Sometimes the Model doth Preach: Quantifying Religious Bias in Open LLMs through Demographic Analysis in Asian Nations
by: Shankar, Hari, et al.
Published: (2025)
by: Shankar, Hari, et al.
Published: (2025)
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions
by: Morabito, Robert, et al.
Published: (2024)
by: Morabito, Robert, et al.
Published: (2024)
Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
by: Kaur, Navreet, et al.
Published: (2025)
by: Kaur, Navreet, et al.
Published: (2025)
MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platform
by: Jung, Hayoung, et al.
Published: (2025)
by: Jung, Hayoung, et al.
Published: (2025)
ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
Understanding The Effect Of Temperature On Alignment With Human Opinions
by: Pavlovic, Maja, et al.
Published: (2024)
by: Pavlovic, Maja, et al.
Published: (2024)
Towards New Benchmark for AI Alignment & Sentiment Analysis in Socially Important Issues: A Comparative Study of Human and LLMs in the Context of AGI
by: Bojic, Ljubisa, et al.
Published: (2025)
by: Bojic, Ljubisa, et al.
Published: (2025)
The Impact and Opportunities of Generative AI in Fact-Checking
by: Wolfe, Robert, et al.
Published: (2024)
by: Wolfe, Robert, et al.
Published: (2024)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
by: Majumdar, Ayan, et al.
Published: (2025)
by: Majumdar, Ayan, et al.
Published: (2025)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
by: Akpinar, Nil-Jana, et al.
Published: (2026)
by: Akpinar, Nil-Jana, et al.
Published: (2026)
Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
by: Prama, Tabia Tanzin, et al.
Published: (2025)
by: Prama, Tabia Tanzin, et al.
Published: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
by: Cai, Yunna, et al.
Published: (2025)
by: Cai, Yunna, et al.
Published: (2025)
Exposure to Content Written by Large Language Models Can Reduce Stigma Around Opioid Use Disorder in Online Communities
by: Mittal, Shravika, et al.
Published: (2025)
by: Mittal, Shravika, et al.
Published: (2025)
Evaluating the Simulation of Human Personality-Driven Susceptibility to Misinformation with LLMs
by: Pratelli, Manuel, et al.
Published: (2025)
by: Pratelli, Manuel, et al.
Published: (2025)
Measuring Fine-Grained Negotiation Tactics of Humans and LLMs in Diplomacy
by: Li, Wenkai, et al.
Published: (2025)
by: Li, Wenkai, et al.
Published: (2025)
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data
by: Mori, Shinka, et al.
Published: (2024)
by: Mori, Shinka, et al.
Published: (2024)
Epistemic Alignment: A Mediating Framework for User-LLM Knowledge Delivery
by: Clark, Nicholas, et al.
Published: (2025)
by: Clark, Nicholas, et al.
Published: (2025)
Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias
by: Tonneau, Manuel, et al.
Published: (2026)
by: Tonneau, Manuel, et al.
Published: (2026)
Expressing Social Emotions: Misalignment Between LLMs and Human Cultural Emotion Norms
by: Bhattacharyya, Sree, et al.
Published: (2026)
by: Bhattacharyya, Sree, et al.
Published: (2026)
When Attention Becomes Exposure in Generative Search
by: Alipour, Shayan, et al.
Published: (2026)
by: Alipour, Shayan, et al.
Published: (2026)
Investigating Political and Demographic Associations in Large Language Models Through Moral Foundations Theory
by: Smith-Vaniz, Nicole, et al.
Published: (2025)
by: Smith-Vaniz, Nicole, et al.
Published: (2025)
Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting
by: Fillies, Jan, et al.
Published: (2025)
by: Fillies, Jan, et al.
Published: (2025)
Similar Items
-
The Gray Area: Characterizing Moderator Disagreement on Reddit
by: Alipour, Shayan, et al.
Published: (2026) -
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
by: Sen, Indira, et al.
Published: (2023) -
The Unseen Targets of Hate -- A Systematic Review of Hateful Communication Datasets
by: Yu, Zehui, et al.
Published: (2024) -
The Implications of Open Generative Models in Human-Centered Data Science Work: A Case Study with Fact-Checking Organizations
by: Wolfe, Robert, et al.
Published: (2024) -
A Multilingual Similarity Dataset for News Article Frame
by: Chen, Xi, et al.
Published: (2024)