Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
Fuente:
arXiv
Guardado en:
| Autores principales: | Chehbouni, Khaoula, Carr, Jonathan Colaço, More, Yash, Cheung, Jackie CK, Farnadi, Golnoosh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
por: Chehbouni, Khaoula, et al.
Publicado: (2024)
por: Chehbouni, Khaoula, et al.
Publicado: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
por: Chehbouni, Khaoula, et al.
Publicado: (2025)
por: Chehbouni, Khaoula, et al.
Publicado: (2025)
Fairness in Federated Learning: Fairness for Whom?
por: Taik, Afaf, et al.
Publicado: (2025)
por: Taik, Afaf, et al.
Publicado: (2025)
Enhancing Privacy in the Early Detection of Sexual Predators Through Federated Learning and Differential Privacy
por: Chehbouni, Khaoula, et al.
Publicado: (2025)
por: Chehbouni, Khaoula, et al.
Publicado: (2025)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
por: More, Yash, et al.
Publicado: (2024)
por: More, Yash, et al.
Publicado: (2024)
Understanding Intrinsic Socioeconomic Biases in Large Language Models
por: Arzaghi, Mina, et al.
Publicado: (2024)
por: Arzaghi, Mina, et al.
Publicado: (2024)
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
por: Mireshghallah, Niloofar, et al.
Publicado: (2024)
por: Mireshghallah, Niloofar, et al.
Publicado: (2024)
Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
por: Arzaghi', 'Mina, et al.
Publicado: (2025)
por: Arzaghi', 'Mina, et al.
Publicado: (2025)
Auditing Agent Harness Safety
por: Liu, Chengzhi, et al.
Publicado: (2026)
por: Liu, Chengzhi, et al.
Publicado: (2026)
LoRA Provides Differential Privacy by Design via Random Sketching
por: Malekmohammadi, Saber, et al.
Publicado: (2024)
por: Malekmohammadi, Saber, et al.
Publicado: (2024)
Multilingual Hallucination Gaps in Large Language Models
por: Chataigner, Cléa, et al.
Publicado: (2024)
por: Chataigner, Cléa, et al.
Publicado: (2024)
Crossing Boundaries: Leveraging Semantic Divergences to Explore Cultural Novelty in Cooking Recipes
por: Carichon, Florian, et al.
Publicado: (2025)
por: Carichon, Florian, et al.
Publicado: (2025)
Causal Fair Metric: Bridging Causality, Individual Fairness, and Adversarial Robustness
por: Ehyaei, Ahmad-Reza, et al.
Publicado: (2023)
por: Ehyaei, Ahmad-Reza, et al.
Publicado: (2023)
Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language Models
por: Zhu, Zhaowei, et al.
Publicado: (2023)
por: Zhu, Zhaowei, et al.
Publicado: (2023)
Position: Cracking the Code of Cascading Disparity Towards Marginalized Communities
por: Farnadi, Golnoosh, et al.
Publicado: (2024)
por: Farnadi, Golnoosh, et al.
Publicado: (2024)
Dishonesty in Helpful and Harmless Alignment
por: Huang, Youcheng, et al.
Publicado: (2024)
por: Huang, Youcheng, et al.
Publicado: (2024)
GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews
por: Darrin, Maxime, et al.
Publicado: (2024)
por: Darrin, Maxime, et al.
Publicado: (2024)
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
por: Vaugrante, Laurène, et al.
Publicado: (2025)
por: Vaugrante, Laurène, et al.
Publicado: (2025)
Say It Another Way: Auditing LLMs with a User-Grounded Automated Paraphrasing Framework
por: Chataigner, Cléa, et al.
Publicado: (2025)
por: Chataigner, Cléa, et al.
Publicado: (2025)
The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
por: Carichon, Florian, et al.
Publicado: (2025)
por: Carichon, Florian, et al.
Publicado: (2025)
Beyond Accuracy: Diagnosing Algebraic Reasoning Failures in LLMs Across Nine Complexity Dimensions
por: Patil, Parth, et al.
Publicado: (2026)
por: Patil, Parth, et al.
Publicado: (2026)
On the Suitability of LLM-Driven Agents for Dark Pattern Audits
por: Sun, Chen, et al.
Publicado: (2026)
por: Sun, Chen, et al.
Publicado: (2026)
Embedding Cultural Diversity in Prototype-based Recommender Systems
por: Moradi, Armin, et al.
Publicado: (2024)
por: Moradi, Armin, et al.
Publicado: (2024)
Different Horses for Different Courses: Comparing Bias Mitigation Algorithms in ML
por: Ganesh, Prakhar, et al.
Publicado: (2024)
por: Ganesh, Prakhar, et al.
Publicado: (2024)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
por: Nghiem, Huy, et al.
Publicado: (2025)
por: Nghiem, Huy, et al.
Publicado: (2025)
Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles
por: Huang, Yue, et al.
Publicado: (2025)
por: Huang, Yue, et al.
Publicado: (2025)
Fairness Incentives in Response to Unfair Dynamic Pricing
por: Thibodeau, Jesse, et al.
Publicado: (2024)
por: Thibodeau, Jesse, et al.
Publicado: (2024)
Auditing Stance Asymmetry in Generative Explanations
por: Han, Jiarui
Publicado: (2026)
por: Han, Jiarui
Publicado: (2026)
The Cost of Arbitrariness for Individuals: Examining the Legal and Technical Challenges of Model Multiplicity
por: Ganesh, Prakhar, et al.
Publicado: (2024)
por: Ganesh, Prakhar, et al.
Publicado: (2024)
Promoting Fair Vaccination Strategies Through Influence Maximization: A Case Study on COVID-19 Spread
por: Neophytou, Nicola, et al.
Publicado: (2024)
por: Neophytou, Nicola, et al.
Publicado: (2024)
Too Helpful, Too Harmless, Too Honest or Just Right?
por: Kashyap, Gautam Siddharth, et al.
Publicado: (2025)
por: Kashyap, Gautam Siddharth, et al.
Publicado: (2025)
Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
por: Stamatis, Caitlin A., et al.
Publicado: (2026)
por: Stamatis, Caitlin A., et al.
Publicado: (2026)
AuditWen:An Open-Source Large Language Model for Audit
por: Huang, Jiajia, et al.
Publicado: (2024)
por: Huang, Jiajia, et al.
Publicado: (2024)
Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
por: Kassem, Aly M., et al.
Publicado: (2025)
por: Kassem, Aly M., et al.
Publicado: (2025)
The Silicon Ceiling: Auditing GPT's Race and Gender Biases in Hiring
por: Armstrong, Lena, et al.
Publicado: (2024)
por: Armstrong, Lena, et al.
Publicado: (2024)
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
por: Arnaiz-Rodriguez, Adrian, et al.
Publicado: (2025)
por: Arnaiz-Rodriguez, Adrian, et al.
Publicado: (2025)
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
por: Hartmann, David, et al.
Publicado: (2026)
por: Hartmann, David, et al.
Publicado: (2026)
Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor
por: Baluja, Ashwin
Publicado: (2024)
por: Baluja, Ashwin
Publicado: (2024)
Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
por: Dutta, Arka, et al.
Publicado: (2023)
por: Dutta, Arka, et al.
Publicado: (2023)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
por: Li, Jing-Jing, et al.
Publicado: (2024)
por: Li, Jing-Jing, et al.
Publicado: (2024)
Ejemplares similares
-
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
por: Chehbouni, Khaoula, et al.
Publicado: (2024) -
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
por: Chehbouni, Khaoula, et al.
Publicado: (2025) -
Fairness in Federated Learning: Fairness for Whom?
por: Taik, Afaf, et al.
Publicado: (2025) -
Enhancing Privacy in the Early Detection of Sexual Predators Through Federated Learning and Differential Privacy
por: Chehbouni, Khaoula, et al.
Publicado: (2025) -
Towards More Realistic Extraction Attacks: An Adversarial Perspective
por: More, Yash, et al.
Publicado: (2024)