EigenBench: A Comparative Behavioral Measure of Value Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Jonathn, Piff, Leonhard, Sana, Suvadip, Li, Jasmine X., Levine, Lionel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026)
by: Lee, Jaehyeok, et al.
Published: (2026)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026)
by: Gringras, David
Published: (2026)
REQUAL-LM: Reliability and Equity through Aggregation in Large Language Models
by: Ebrahimi, Sana, et al.
Published: (2024)
by: Ebrahimi, Sana, et al.
Published: (2024)
ProgressGym: Alignment with a Millennium of Moral Progress
by: Qiu, Tianyi, et al.
Published: (2024)
by: Qiu, Tianyi, et al.
Published: (2024)
Do Large Language Models Walk Their Talk? Measuring the Gap Between Implicit Associations, Self-Report, and Behavioral Altruism
by: Andric, Sandro
Published: (2025)
by: Andric, Sandro
Published: (2025)
AXOLOTL: Fairness through Assisted Self-Debiasing of Large Language Model Outputs
by: Ebrahimi, Sana, et al.
Published: (2024)
by: Ebrahimi, Sana, et al.
Published: (2024)
HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
by: Liu, Haokun, et al.
Published: (2025)
by: Liu, Haokun, et al.
Published: (2025)
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
by: Huang, Saffron, et al.
Published: (2025)
by: Huang, Saffron, et al.
Published: (2025)
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
by: Chiu, Yu Ying, et al.
Published: (2025)
by: Chiu, Yu Ying, et al.
Published: (2025)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Deliberative Alignment: Reasoning Enables Safer Language Models
by: Guan, Melody Y., et al.
Published: (2024)
by: Guan, Melody Y., et al.
Published: (2024)
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
Reward Models Inherit Value Biases from Pretraining
by: Christian, Brian, et al.
Published: (2026)
by: Christian, Brian, et al.
Published: (2026)
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
by: Anwar, Usman, et al.
Published: (2024)
by: Anwar, Usman, et al.
Published: (2024)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
Does Cross-Cultural Alignment Change the Commonsense Morality of Language Models?
by: Jinnai, Yuu
Published: (2024)
by: Jinnai, Yuu
Published: (2024)
Prompt-Counterfactual Explanations for Generative AI System Behavior
by: Goethals, Sofie, et al.
Published: (2026)
by: Goethals, Sofie, et al.
Published: (2026)
Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models
by: Basu, Abhinaba, et al.
Published: (2026)
by: Basu, Abhinaba, et al.
Published: (2026)
GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning
by: Yu, Jeffy, et al.
Published: (2024)
by: Yu, Jeffy, et al.
Published: (2024)
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
by: Han, Pengrui, et al.
Published: (2025)
by: Han, Pengrui, et al.
Published: (2025)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
by: Chiu, Yu Ying, et al.
Published: (2025)
by: Chiu, Yu Ying, et al.
Published: (2025)
Surveying Attitudinal Alignment Between Large Language Models Vs. Humans Towards 17 Sustainable Development Goals
by: Wu, Qingyang, et al.
Published: (2024)
by: Wu, Qingyang, et al.
Published: (2024)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
by: Chien, Jennifer, et al.
Published: (2024)
by: Chien, Jennifer, et al.
Published: (2024)
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
by: Arakelyan, Erik, et al.
Published: (2024)
by: Arakelyan, Erik, et al.
Published: (2024)
Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation
by: Franceschelli, Giorgio, et al.
Published: (2025)
by: Franceschelli, Giorgio, et al.
Published: (2025)
From Perceptions to Decisions: Wildfire Evacuation Decision Prediction with Behavioral Theory-informed LLMs
by: Chen, Ruxiao, et al.
Published: (2025)
by: Chen, Ruxiao, et al.
Published: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
CEGI: Measuring the trade-off between efficiency and carbon emissions for SLMs and VLMs
by: Kumar, Abhas, et al.
Published: (2024)
by: Kumar, Abhas, et al.
Published: (2024)
Computational Measurement of Political Positions: A Review of Text-Based Ideal Point Estimation Algorithms
by: Parschan, Patrick, et al.
Published: (2025)
by: Parschan, Patrick, et al.
Published: (2025)
LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models
by: Meadows, Gwenyth Isobel, et al.
Published: (2024)
by: Meadows, Gwenyth Isobel, et al.
Published: (2024)
M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation
by: Jeong, Jihoon
Published: (2026)
by: Jeong, Jihoon
Published: (2026)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
From Data to Behavior: Predicting Unintended Model Behaviors Before Training
by: Wang, Mengru, et al.
Published: (2026)
by: Wang, Mengru, et al.
Published: (2026)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
by: Sturgeon, Benjamin, et al.
Published: (2025)
by: Sturgeon, Benjamin, et al.
Published: (2025)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
by: Ghaboura, Sara, et al.
Published: (2024)
by: Ghaboura, Sara, et al.
Published: (2024)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
Similar Items
-
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
by: Xiao, Yuxin, et al.
Published: (2025) -
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026) -
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025) -
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026) -
REQUAL-LM: Reliability and Equity through Aggregation in Large Language Models
by: Ebrahimi, Sana, et al.
Published: (2024)