SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jing-Jing, Pyatkin, Valentina, Kleiman-Weiner, Max, Jiang, Liwei, Dziri, Nouha, Collins, Anne G. E., Borg, Jana Schaich, Sap, Maarten, Choi, Yejin, Levine, Sydney |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
by: Sorensen, Taylor, et al.
Published: (2023)
by: Sorensen, Taylor, et al.
Published: (2023)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024)
by: Han, Seungju, et al.
Published: (2024)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
by: Rao, Kavel, et al.
Published: (2023)
by: Rao, Kavel, et al.
Published: (2023)
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
by: Graf, Victoria, et al.
Published: (2026)
by: Graf, Victoria, et al.
Published: (2026)
PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
by: Kumar, Priyanshu, et al.
Published: (2025)
by: Kumar, Priyanshu, et al.
Published: (2025)
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
by: Jiang, Liwei, et al.
Published: (2025)
by: Jiang, Liwei, et al.
Published: (2025)
Why (not) use AI? Analyzing People's Reasoning and Conditions for AI Acceptability
by: Mun, Jimin, et al.
Published: (2025)
by: Mun, Jimin, et al.
Published: (2025)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
by: Bennion, Jonathan, et al.
Published: (2025)
by: Bennion, Jonathan, et al.
Published: (2025)
What Is Required for Empathic AI? It Depends, and Why That Matters for AI Developers and Users
by: Borg, Jana Schaich, et al.
Published: (2024)
by: Borg, Jana Schaich, et al.
Published: (2024)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
by: Qiu, Linlu, et al.
Published: (2023)
by: Qiu, Linlu, et al.
Published: (2023)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
by: Zhou, Kaitlyn, et al.
Published: (2024)
by: Zhou, Kaitlyn, et al.
Published: (2024)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
by: Jiang, Liwei, et al.
Published: (2024)
by: Jiang, Liwei, et al.
Published: (2024)
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
by: Baheti, Ashutosh, et al.
Published: (2024)
by: Baheti, Ashutosh, et al.
Published: (2024)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
by: Lin, Bill Yuchen, et al.
Published: (2024)
by: Lin, Bill Yuchen, et al.
Published: (2024)
Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics
by: Liu, Jiarui, et al.
Published: (2025)
by: Liu, Jiarui, et al.
Published: (2025)
Value Internalization: Learning and Generalizing from Social Reward
by: Rong, Frieda, et al.
Published: (2024)
by: Rong, Frieda, et al.
Published: (2024)
Evaluating LLMs in Open-Source Games
by: Sistla, Swadesh, et al.
Published: (2025)
by: Sistla, Swadesh, et al.
Published: (2025)
Can Language Models Reason about Individualistic Human Values and Preferences?
by: Jiang, Liwei, et al.
Published: (2024)
by: Jiang, Liwei, et al.
Published: (2024)
When Should AI Read the Room? Public Perceptions of Social Intelligence in AI Agents
by: Mathur, Leena, et al.
Published: (2026)
by: Mathur, Leena, et al.
Published: (2026)
HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions
by: Zhou, Xuhui, et al.
Published: (2024)
by: Zhou, Xuhui, et al.
Published: (2024)
RewardBench: Evaluating Reward Models for Language Modeling
by: Lambert, Nathan, et al.
Published: (2024)
by: Lambert, Nathan, et al.
Published: (2024)
Boundedly Rational Meta-Learning in Sequential Consumer Choice
by: Khosravi, Mehrzad, et al.
Published: (2026)
by: Khosravi, Mehrzad, et al.
Published: (2026)
Estimating the Empowerment of Language Model Agents
by: Song, Jinyeop, et al.
Published: (2025)
by: Song, Jinyeop, et al.
Published: (2025)
When Empowerment Disempowers
by: Yang, Claire, et al.
Published: (2025)
by: Yang, Claire, et al.
Published: (2025)
Intuitions of Compromise: Utilitarianism vs. Contractualism
by: Moore, Jared, et al.
Published: (2024)
by: Moore, Jared, et al.
Published: (2024)
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
by: Ghate, Kshitish, et al.
Published: (2025)
by: Ghate, Kshitish, et al.
Published: (2025)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
by: Mendelsohn, Julia, et al.
Published: (2023)
by: Mendelsohn, Julia, et al.
Published: (2023)
Language Model Alignment in Multilingual Trolley Problems
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
The Art of Saying No: Contextual Noncompliance in Language Models
by: Brahman, Faeze, et al.
Published: (2024)
by: Brahman, Faeze, et al.
Published: (2024)
When Is It Acceptable to Break the Rules? Knowledge Representation of Moral Judgement Based on Empirical Data
by: Awad, Edmond, et al.
Published: (2022)
by: Awad, Edmond, et al.
Published: (2022)
Preserving Sense of Agency: User Preferences for Robot Autonomy and User Control across Household Tasks
by: Yang, Claire, et al.
Published: (2025)
by: Yang, Claire, et al.
Published: (2025)
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
by: Sorensen, Taylor, et al.
Published: (2025)
by: Sorensen, Taylor, et al.
Published: (2025)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
by: Lu, Ximing, et al.
Published: (2024)
by: Lu, Ximing, et al.
Published: (2024)
CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
by: Li, Huihan, et al.
Published: (2024)
by: Li, Huihan, et al.
Published: (2024)
On the Pros and Cons of Active Learning for Moral Preference Elicitation
by: Keswani, Vijay, et al.
Published: (2024)
by: Keswani, Vijay, et al.
Published: (2024)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
by: Zheng, Mingqian, et al.
Published: (2026)
by: Zheng, Mingqian, et al.
Published: (2026)
Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
by: Mun, Jimin, et al.
Published: (2024)
by: Mun, Jimin, et al.
Published: (2024)
The Lock-in Hypothesis: Stagnation by Algorithm
by: Qiu, Tianyi Alex, et al.
Published: (2025)
by: Qiu, Tianyi Alex, et al.
Published: (2025)
Similar Items
-
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026) -
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
by: Sorensen, Taylor, et al.
Published: (2023) -
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024) -
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
by: Vijayvargiya, Sanidhya, et al.
Published: (2025) -
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
by: Rao, Kavel, et al.
Published: (2023)