Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Ashkinaze, Joshua, Shen, Hua, Avula, Saipranav, Gilbert, Eric, Budak, Ceren |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment
by: Ashkinaze, Joshua, et al.
Published: (2024)
by: Ashkinaze, Joshua, et al.
Published: (2024)
Plurals: A System for Guiding LLMs Via Simulated Social Ensembles
by: Ashkinaze, Joshua, et al.
Published: (2024)
by: Ashkinaze, Joshua, et al.
Published: (2024)
Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms
by: Ashkinaze, Joshua, et al.
Published: (2024)
by: Ashkinaze, Joshua, et al.
Published: (2024)
The Dynamics of (Not) Unfollowing Misinformation Spreaders
by: Ashkinaze, Joshua, et al.
Published: (2024)
by: Ashkinaze, Joshua, et al.
Published: (2024)
When People are Floods: Analyzing Dehumanizing Metaphors in Immigration Discourse with Large Language Models
by: Mendelsohn, Julia, et al.
Published: (2025)
by: Mendelsohn, Julia, et al.
Published: (2025)
LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models
by: Meadows, Gwenyth Isobel, et al.
Published: (2024)
by: Meadows, Gwenyth Isobel, et al.
Published: (2024)
Growth First, Care Second? Tracing the Landscape of LLM Value Preferences in Everyday Dilemmas
by: Chen, Zhiyi, et al.
Published: (2026)
by: Chen, Zhiyi, et al.
Published: (2026)
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
by: Yao, Jing, et al.
Published: (2025)
by: Yao, Jing, et al.
Published: (2025)
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
by: Zhang, Yazhou, et al.
Published: (2025)
by: Zhang, Yazhou, et al.
Published: (2025)
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
by: Jiang, Han, et al.
Published: (2024)
by: Jiang, Han, et al.
Published: (2024)
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
by: Huang, Saffron, et al.
Published: (2025)
by: Huang, Saffron, et al.
Published: (2025)
EigenBench: A Comparative Behavioral Measure of Value Alignment
by: Chang, Jonathn, et al.
Published: (2025)
by: Chang, Jonathn, et al.
Published: (2025)
ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models
by: Zhang, Zhaowei, et al.
Published: (2023)
by: Zhang, Zhaowei, et al.
Published: (2023)
Measuring Political Preferences in AI Systems: An Integrative Approach
by: Rozado, David
Published: (2025)
by: Rozado, David
Published: (2025)
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
by: Duan, Shitong, et al.
Published: (2023)
by: Duan, Shitong, et al.
Published: (2023)
Reward Models Inherit Value Biases from Pretraining
by: Christian, Brian, et al.
Published: (2026)
by: Christian, Brian, et al.
Published: (2026)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026)
by: Lee, Jaehyeok, et al.
Published: (2026)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
by: Jiang, Han, et al.
Published: (2025)
by: Jiang, Han, et al.
Published: (2025)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
by: Qu, Jinxian, et al.
Published: (2026)
by: Qu, Jinxian, et al.
Published: (2026)
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts
by: Zhang, Wenjing, et al.
Published: (2025)
by: Zhang, Wenjing, et al.
Published: (2025)
Prompt Design Matters for Computational Social Science Tasks but in Unpredictable Ways
by: Atreja, Shubham, et al.
Published: (2024)
by: Atreja, Shubham, et al.
Published: (2024)
When AI Speaks, Whose Values Does It Express? A Cross-Cultural Audit of Individualism-Collectivism Bias in Large Language Models
by: Venkata, Pruthvinath Jeripity
Published: (2026)
by: Venkata, Pruthvinath Jeripity
Published: (2026)
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
by: Greco, Candida M., et al.
Published: (2026)
by: Greco, Candida M., et al.
Published: (2026)
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
by: Qiu, Peiran, et al.
Published: (2025)
by: Qiu, Peiran, et al.
Published: (2025)
ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
Temporal Preferences in Language Models for Long-Horizon Assistance
by: Mazyaki, Ali, et al.
Published: (2025)
by: Mazyaki, Ali, et al.
Published: (2025)
DeepTutor: Towards Agentic Personalized Tutoring
by: Zhao, Bingxi, et al.
Published: (2026)
by: Zhao, Bingxi, et al.
Published: (2026)
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)
by: Fan, Xianzhe, et al.
Published: (2024)
neuralFOMO: Can LLMs Handle Being Second Best? Measuring Envy-Like Preferences in Multi-Agent Settings
by: Ramamoorthy, Arnav, et al.
Published: (2025)
by: Ramamoorthy, Arnav, et al.
Published: (2025)
GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning
by: Yu, Jeffy, et al.
Published: (2024)
by: Yu, Jeffy, et al.
Published: (2024)
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
by: Shen, Hua, et al.
Published: (2025)
by: Shen, Hua, et al.
Published: (2025)
ShaRP: Explaining Rankings and Preferences with Shapley Values
by: Pliatsika, Venetia, et al.
Published: (2024)
by: Pliatsika, Venetia, et al.
Published: (2024)
The Political Preferences of LLMs
by: Rozado, David
Published: (2024)
by: Rozado, David
Published: (2024)
Critical Foreign Policy Decisions (CFPD)-Benchmark: Measuring Diplomatic Preferences in Large Language Models
by: Jensen, Benjamin, et al.
Published: (2025)
by: Jensen, Benjamin, et al.
Published: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation
by: Franceschelli, Giorgio, et al.
Published: (2025)
by: Franceschelli, Giorgio, et al.
Published: (2025)
Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions
by: Rakshit, Sushrita, et al.
Published: (2026)
by: Rakshit, Sushrita, et al.
Published: (2026)
Measuring Human Contribution in AI-Assisted Content Generation
by: Xie, Yueqi, et al.
Published: (2024)
by: Xie, Yueqi, et al.
Published: (2024)
Emotion-Aware Embedding Fusion in LLMs (Flan-T5, LLAMA 2, DeepSeek-R1, and ChatGPT 4) for Intelligent Response Generation
by: Rasool, Abdur, et al.
Published: (2024)
by: Rasool, Abdur, et al.
Published: (2024)
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
by: Zhang, Guo-Biao, et al.
Published: (2026)
by: Zhang, Guo-Biao, et al.
Published: (2026)
Similar Items
-
How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment
by: Ashkinaze, Joshua, et al.
Published: (2024) -
Plurals: A System for Guiding LLMs Via Simulated Social Ensembles
by: Ashkinaze, Joshua, et al.
Published: (2024) -
Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms
by: Ashkinaze, Joshua, et al.
Published: (2024) -
The Dynamics of (Not) Unfollowing Misinformation Spreaders
by: Ashkinaze, Joshua, et al.
Published: (2024) -
When People are Floods: Analyzing Dehumanizing Metaphors in Immigration Discourse with Large Language Models
by: Mendelsohn, Julia, et al.
Published: (2025)