From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
Fuente:
arXiv
Saved in:
| Main Authors: | Chehbouni, Khaoula, Roshan, Megha, Ma, Emmanuel, Wei, Futian Andrew, Taik, Afaf, Cheung, Jackie CK, Farnadi, Golnoosh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fairness in Federated Learning: Fairness for Whom?
by: Taik, Afaf, et al.
Published: (2025)
by: Taik, Afaf, et al.
Published: (2025)
Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
by: Chehbouni, Khaoula, et al.
Published: (2024)
by: Chehbouni, Khaoula, et al.
Published: (2024)
Enhancing Privacy in the Early Detection of Sexual Predators Through Federated Learning and Differential Privacy
by: Chehbouni, Khaoula, et al.
Published: (2025)
by: Chehbouni, Khaoula, et al.
Published: (2025)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
by: Chehbouni, Khaoula, et al.
Published: (2025)
by: Chehbouni, Khaoula, et al.
Published: (2025)
Multilingual Hallucination Gaps in Large Language Models
by: Chataigner, Cléa, et al.
Published: (2024)
by: Chataigner, Cléa, et al.
Published: (2024)
Promoting Fair Vaccination Strategies Through Influence Maximization: A Case Study on COVID-19 Spread
by: Neophytou, Nicola, et al.
Published: (2024)
by: Neophytou, Nicola, et al.
Published: (2024)
Differentially Private Clustered Federated Learning
by: Malekmohammadi, Saber, et al.
Published: (2024)
by: Malekmohammadi, Saber, et al.
Published: (2024)
Systemizing Multiplicity: The Curious Case of Arbitrariness in Machine Learning
by: Ganesh, Prakhar, et al.
Published: (2025)
by: Ganesh, Prakhar, et al.
Published: (2025)
Fairness Incentives in Response to Unfair Dynamic Pricing
by: Thibodeau, Jesse, et al.
Published: (2024)
by: Thibodeau, Jesse, et al.
Published: (2024)
Say It Another Way: Auditing LLMs with a User-Grounded Automated Paraphrasing Framework
by: Chataigner, Cléa, et al.
Published: (2025)
by: Chataigner, Cléa, et al.
Published: (2025)
Balancing Profit and Fairness in Risk-Based Pricing Markets
by: Thibodeau, Jesse, et al.
Published: (2025)
by: Thibodeau, Jesse, et al.
Published: (2025)
LoRA Provides Differential Privacy by Design via Random Sketching
by: Malekmohammadi, Saber, et al.
Published: (2024)
by: Malekmohammadi, Saber, et al.
Published: (2024)
Understanding Intrinsic Socioeconomic Biases in Large Language Models
by: Arzaghi, Mina, et al.
Published: (2024)
by: Arzaghi, Mina, et al.
Published: (2024)
Crossing Boundaries: Leveraging Semantic Divergences to Explore Cultural Novelty in Cooking Recipes
by: Carichon, Florian, et al.
Published: (2025)
by: Carichon, Florian, et al.
Published: (2025)
GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews
by: Darrin, Maxime, et al.
Published: (2024)
by: Darrin, Maxime, et al.
Published: (2024)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026)
by: Kaunismaa, Jackson, et al.
Published: (2026)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
by: More, Yash, et al.
Published: (2024)
by: More, Yash, et al.
Published: (2024)
Position: Cracking the Code of Cascading Disparity Towards Marginalized Communities
by: Farnadi, Golnoosh, et al.
Published: (2024)
by: Farnadi, Golnoosh, et al.
Published: (2024)
A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
by: Harvey, Emma, et al.
Published: (2025)
by: Harvey, Emma, et al.
Published: (2025)
AI Safety Training Can be Clinically Harmful
by: BN, Suhas, et al.
Published: (2026)
by: BN, Suhas, et al.
Published: (2026)
Causal Fair Metric: Bridging Causality, Individual Fairness, and Adversarial Robustness
by: Ehyaei, Ahmad-Reza, et al.
Published: (2023)
by: Ehyaei, Ahmad-Reza, et al.
Published: (2023)
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025)
by: Corvi, Emily, et al.
Published: (2025)
Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
by: Arzaghi', 'Mina, et al.
Published: (2025)
by: Arzaghi', 'Mina, et al.
Published: (2025)
The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
by: Carichon, Florian, et al.
Published: (2025)
by: Carichon, Florian, et al.
Published: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
Residual Connections Harm Generative Representation Learning
by: Zhang, Xiao, et al.
Published: (2024)
by: Zhang, Xiao, et al.
Published: (2024)
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
by: Yi, Biao, et al.
Published: (2025)
by: Yi, Biao, et al.
Published: (2025)
Gendered Inequalities in Online Harms: Fear, Safety Work, and Online Participation
by: Enock, Florence E., et al.
Published: (2024)
by: Enock, Florence E., et al.
Published: (2024)
Embedding Cultural Diversity in Prototype-based Recommender Systems
by: Moradi, Armin, et al.
Published: (2024)
by: Moradi, Armin, et al.
Published: (2024)
Different Horses for Different Courses: Comparing Bias Mitigation Algorithms in ML
by: Ganesh, Prakhar, et al.
Published: (2024)
by: Ganesh, Prakhar, et al.
Published: (2024)
Vernacularizing Taxonomies of Harm is Essential for Operationalizing Holistic AI Safety
by: Kennedy, Wm. Matthew, et al.
Published: (2024)
by: Kennedy, Wm. Matthew, et al.
Published: (2024)
Should LLM Safety Be More Than Refusing Harmful Instructions?
by: Maskey, Utsav, et al.
Published: (2025)
by: Maskey, Utsav, et al.
Published: (2025)
More of the Same: Persistent Representational Harms Under Increased Representation
by: Mickel, Jennifer, et al.
Published: (2025)
by: Mickel, Jennifer, et al.
Published: (2025)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
by: Kassem, Aly M., et al.
Published: (2025)
by: Kassem, Aly M., et al.
Published: (2025)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
by: Chien, Jennifer, et al.
Published: (2024)
by: Chien, Jennifer, et al.
Published: (2024)
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
by: Mireshghallah, Niloofar, et al.
Published: (2024)
by: Mireshghallah, Niloofar, et al.
Published: (2024)
Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators
by: Hutiri, Wiebke, et al.
Published: (2024)
by: Hutiri, Wiebke, et al.
Published: (2024)
Similar Items
-
Fairness in Federated Learning: Fairness for Whom?
by: Taik, Afaf, et al.
Published: (2025) -
Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
by: Chehbouni, Khaoula, et al.
Published: (2024) -
Enhancing Privacy in the Early Detection of Sexual Predators Through Federated Learning and Differential Privacy
by: Chehbouni, Khaoula, et al.
Published: (2025) -
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
by: Chehbouni, Khaoula, et al.
Published: (2025) -
Multilingual Hallucination Gaps in Large Language Models
by: Chataigner, Cléa, et al.
Published: (2024)