Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Saffron, Durmus, Esin, McCain, Miles, Handa, Kunal, Tamkin, Alex, Hong, Jerry, Stern, Michael, Somani, Arushi, Zhang, Xiuruo, Ganguli, Deep
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915252439875584
author Huang, Saffron
Durmus, Esin
McCain, Miles
Handa, Kunal
Tamkin, Alex
Hong, Jerry
Stern, Michael
Somani, Arushi
Zhang, Xiuruo
Ganguli, Deep
author_facet Huang, Saffron
Durmus, Esin
McCain, Miles
Handa, Kunal
Tamkin, Alex
Hong, Jerry
Stern, Michael
Somani, Arushi
Zhang, Xiuruo
Ganguli, Deep
contents AI assistants can impart value judgments that shape people's decisions and worldviews, yet little is known empirically about what values these systems rely on in practice. To address this, we develop a bottom-up, privacy-preserving method to extract the values (normative considerations stated or demonstrated in model responses) that Claude 3 and 3.5 models exhibit in hundreds of thousands of real-world interactions. We empirically discover and taxonomize 3,307 AI values and study how they vary by context. We find that Claude expresses many practical and epistemic values, and typically supports prosocial human values while resisting values like "moral nihilism". While some values appear consistently across contexts (e.g. "transparency"), many are more specialized and context-dependent, reflecting the diversity of human interlocutors and their varied contexts. For example, "harm prevention" emerges when Claude resists users, "historical accuracy" when responding to queries about controversial events, "healthy boundaries" when asked for relationship advice, and "human agency" in technology ethics discussions. By providing the first large-scale empirical mapping of AI values in deployment, our work creates a foundation for more grounded evaluation and design of values in AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15236
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
Huang, Saffron
Durmus, Esin
McCain, Miles
Handa, Kunal
Tamkin, Alex
Hong, Jerry
Stern, Michael
Somani, Arushi
Zhang, Xiuruo
Ganguli, Deep
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
AI assistants can impart value judgments that shape people's decisions and worldviews, yet little is known empirically about what values these systems rely on in practice. To address this, we develop a bottom-up, privacy-preserving method to extract the values (normative considerations stated or demonstrated in model responses) that Claude 3 and 3.5 models exhibit in hundreds of thousands of real-world interactions. We empirically discover and taxonomize 3,307 AI values and study how they vary by context. We find that Claude expresses many practical and epistemic values, and typically supports prosocial human values while resisting values like "moral nihilism". While some values appear consistently across contexts (e.g. "transparency"), many are more specialized and context-dependent, reflecting the diversity of human interlocutors and their varied contexts. For example, "harm prevention" emerges when Claude resists users, "historical accuracy" when responding to queries about controversial events, "healthy boundaries" when asked for relationship advice, and "human agency" in technology ethics discussions. By providing the first large-scale empirical mapping of AI values in deployment, our work creates a foundation for more grounded evaluation and design of values in AI systems.
title Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2504.15236