Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cyberey, Hannah, Ji, Yangfeng, Evans, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unsupervised Concept Vector Extraction for Bias Control in LLMs
von: Cyberey, Hannah, et al.
Veröffentlicht: (2025)
von: Cyberey, Hannah, et al.
Veröffentlicht: (2025)
White-Box Sensitivity Auditing with Steering Vectors
von: Cyberey, Hannah, et al.
Veröffentlicht: (2026)
von: Cyberey, Hannah, et al.
Veröffentlicht: (2026)
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
von: Cyberey, Hannah, et al.
Veröffentlicht: (2025)
von: Cyberey, Hannah, et al.
Veröffentlicht: (2025)
Addressing Both Statistical and Causal Gender Fairness in NLP Models
von: Chen, Hannah, et al.
Veröffentlicht: (2024)
von: Chen, Hannah, et al.
Veröffentlicht: (2024)
A Capabilities Approach to Studying Bias and Harm in Language Technologies
von: Nigatu, Hellina Hailu, et al.
Veröffentlicht: (2024)
von: Nigatu, Hellina Hailu, et al.
Veröffentlicht: (2024)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
von: Dorn, Rebecca, et al.
Veröffentlicht: (2024)
von: Dorn, Rebecca, et al.
Veröffentlicht: (2024)
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
von: Arnaiz-Rodriguez, Adrian, et al.
Veröffentlicht: (2025)
von: Arnaiz-Rodriguez, Adrian, et al.
Veröffentlicht: (2025)
Diverse, but Divisive: LLMs Can Exaggerate Gender Differences in Opinion Related to Harms of Misinformation
von: Neumann, Terrence, et al.
Veröffentlicht: (2024)
von: Neumann, Terrence, et al.
Veröffentlicht: (2024)
Taxonomizing Representational Harms using Speech Act Theory
von: Corvi, Emily, et al.
Veröffentlicht: (2025)
von: Corvi, Emily, et al.
Veröffentlicht: (2025)
HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
von: Weissburg, Iain, et al.
Veröffentlicht: (2024)
von: Weissburg, Iain, et al.
Veröffentlicht: (2024)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
von: Mohamadi, Alireza, et al.
Veröffentlicht: (2025)
von: Mohamadi, Alireza, et al.
Veröffentlicht: (2025)
First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows
von: Xing, Sihao, et al.
Veröffentlicht: (2026)
von: Xing, Sihao, et al.
Veröffentlicht: (2026)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
Careless Whisper: Speech-to-Text Hallucination Harms
von: Koenecke, Allison, et al.
Veröffentlicht: (2024)
von: Koenecke, Allison, et al.
Veröffentlicht: (2024)
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
von: Potter, Yujin, et al.
Veröffentlicht: (2024)
von: Potter, Yujin, et al.
Veröffentlicht: (2024)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
von: Choi, Sooyung, et al.
Veröffentlicht: (2025)
von: Choi, Sooyung, et al.
Veröffentlicht: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
LLM-based Semantic Augmentation for Harmful Content Detection
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
von: Ko, Changgeon, et al.
Veröffentlicht: (2024)
von: Ko, Changgeon, et al.
Veröffentlicht: (2024)
Perceived Political Bias in LLMs Reduces Persuasive Abilities
von: DiGiuseppe, Matthew, et al.
Veröffentlicht: (2026)
von: DiGiuseppe, Matthew, et al.
Veröffentlicht: (2026)
EtiCor++: Towards Understanding Etiquettical Bias in LLMs
von: Dwivedi, Ashutosh, et al.
Veröffentlicht: (2025)
von: Dwivedi, Ashutosh, et al.
Veröffentlicht: (2025)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
Questionnaire Responses Do not Capture the Safety of AI Agents
von: Hellrigel-Holderbaum, Max, et al.
Veröffentlicht: (2026)
von: Hellrigel-Holderbaum, Max, et al.
Veröffentlicht: (2026)
PANORAMA: A Dataset and Benchmarks Capturing Decision Trails and Rationales in Patent Examination
von: Lim, Hyunseung, et al.
Veröffentlicht: (2025)
von: Lim, Hyunseung, et al.
Veröffentlicht: (2025)
Sometimes the Model doth Preach: Quantifying Religious Bias in Open LLMs through Demographic Analysis in Asian Nations
von: Shankar, Hari, et al.
Veröffentlicht: (2025)
von: Shankar, Hari, et al.
Veröffentlicht: (2025)
Dual-Metric Evaluation of Social Bias in Large Language Models: Evidence from an Underrepresented Nepali Cultural Context
von: Pandey, Ashish, et al.
Veröffentlicht: (2026)
von: Pandey, Ashish, et al.
Veröffentlicht: (2026)
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs
von: Fernandes, Gustavo Lúcius, et al.
Veröffentlicht: (2026)
von: Fernandes, Gustavo Lúcius, et al.
Veröffentlicht: (2026)
DoDo Learning: DOmain-DemOgraphic Transfer in Language Models for Detecting Abuse Targeted at Public Figures
von: Williams, Angus R., et al.
Veröffentlicht: (2023)
von: Williams, Angus R., et al.
Veröffentlicht: (2023)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text
von: Kanepajs, Arturs, et al.
Veröffentlicht: (2025)
von: Kanepajs, Arturs, et al.
Veröffentlicht: (2025)
Red Lines and Grey Zones in the Fog of War: Benchmarking Legal Risk, Moral Harm, and Regional Bias in Large Language Model Military Decision-Making
von: Drinkall, Toby
Veröffentlicht: (2025)
von: Drinkall, Toby
Veröffentlicht: (2025)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
von: Wachter, Jasmin, et al.
Veröffentlicht: (2025)
von: Wachter, Jasmin, et al.
Veröffentlicht: (2025)
Counterfactual Probing for the Influence of Affect and Specificity on Intergroup Bias
von: Govindarajan, Venkata S, et al.
Veröffentlicht: (2023)
von: Govindarajan, Venkata S, et al.
Veröffentlicht: (2023)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2024)
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2024)
Gender Bias in LLMs: Preliminary Evidence from Shared Parenting Scenario in Czech Family Law
von: Harasta, Jakub, et al.
Veröffentlicht: (2026)
von: Harasta, Jakub, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Unsupervised Concept Vector Extraction for Bias Control in LLMs
von: Cyberey, Hannah, et al.
Veröffentlicht: (2025) -
White-Box Sensitivity Auditing with Steering Vectors
von: Cyberey, Hannah, et al.
Veröffentlicht: (2026) -
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
von: Cyberey, Hannah, et al.
Veröffentlicht: (2025) -
Addressing Both Statistical and Causal Gender Fairness in NLP Models
von: Chen, Hannah, et al.
Veröffentlicht: (2024) -
A Capabilities Approach to Studying Bias and Harm in Language Technologies
von: Nigatu, Hellina Hailu, et al.
Veröffentlicht: (2024)