Generative Value Conflicts Reveal LLM Priorities
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Andy, Ghate, Kshitish, Diab, Mona, Fried, Daniel, Kasirzadeh, Atoosa, Kleiman-Weiner, Max |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Personal Information Parroting in Language Models
by: Subramani, Nishant, et al.
Published: (2026)
by: Subramani, Nishant, et al.
Published: (2026)
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
by: Ghate, Kshitish, et al.
Published: (2025)
by: Ghate, Kshitish, et al.
Published: (2025)
Two Types of AI Existential Risk: Decisive and Accumulative
by: Kasirzadeh, Atoosa
Published: (2024)
by: Kasirzadeh, Atoosa
Published: (2024)
Value Internalization: Learning and Generalizing from Social Reward
by: Rong, Frieda, et al.
Published: (2024)
by: Rong, Frieda, et al.
Published: (2024)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
by: Muhamed, Aashiq, et al.
Published: (2024)
by: Muhamed, Aashiq, et al.
Published: (2024)
Emotion Classification in Low and Moderate Resource Languages
by: Tafreshi, Shabnam, et al.
Published: (2024)
by: Tafreshi, Shabnam, et al.
Published: (2024)
SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
CoRAG: Collaborative Retrieval-Augmented Generation
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Pre-Calc: Learning to Use the Calculator Improves Numeracy in Language Models
by: Veerendranath, Vishruth, et al.
Published: (2024)
by: Veerendranath, Vishruth, et al.
Published: (2024)
The Lock-in Hypothesis: Stagnation by Algorithm
by: Qiu, Tianyi Alex, et al.
Published: (2025)
by: Qiu, Tianyi Alex, et al.
Published: (2025)
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
by: Wedgwood, James, et al.
Published: (2026)
by: Wedgwood, James, et al.
Published: (2026)
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes
by: Ghate, Kshitish, et al.
Published: (2025)
by: Ghate, Kshitish, et al.
Published: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
by: Chopra, Harshita, et al.
Published: (2026)
by: Chopra, Harshita, et al.
Published: (2026)
Evaluating LLMs in Open-Source Games
by: Sistla, Swadesh, et al.
Published: (2025)
by: Sistla, Swadesh, et al.
Published: (2025)
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
by: Chegini, Atoosa, et al.
Published: (2026)
by: Chegini, Atoosa, et al.
Published: (2026)
Evaluating Large Language Model Biases in Persona-Steered Generation
by: Liu, Andy, et al.
Published: (2024)
by: Liu, Andy, et al.
Published: (2024)
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
by: Kazemi, Hamid, et al.
Published: (2026)
by: Kazemi, Hamid, et al.
Published: (2026)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Estimating the Empowerment of Language Model Agents
by: Song, Jinyeop, et al.
Published: (2025)
by: Song, Jinyeop, et al.
Published: (2025)
Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
CLadder: Assessing Causal Reasoning in Language Models
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models
by: Pistilli, Giada, et al.
Published: (2024)
by: Pistilli, Giada, et al.
Published: (2024)
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders
by: Ghate, Kshitish, et al.
Published: (2025)
by: Ghate, Kshitish, et al.
Published: (2025)
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
by: Shan, Zikang, et al.
Published: (2026)
by: Shan, Zikang, et al.
Published: (2026)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
by: Chiu, Yu Ying, et al.
Published: (2024)
by: Chiu, Yu Ying, et al.
Published: (2024)
Explanation Hacking: The perils of algorithmic recourse
by: Sullivan, Emily, et al.
Published: (2024)
by: Sullivan, Emily, et al.
Published: (2024)
Bridging the Gap in the Responsible AI Divides
by: Gyevnár, Bálint, et al.
Published: (2026)
by: Gyevnár, Bálint, et al.
Published: (2026)
Tree Search for Language Model Agents
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
by: Kim, Taesu, et al.
Published: (2024)
by: Kim, Taesu, et al.
Published: (2024)
Generating Pragmatic Examples to Train Neural Program Synthesizers
by: Vaduguru, Saujas, et al.
Published: (2023)
by: Vaduguru, Saujas, et al.
Published: (2023)
REALM: A Dataset of Real-World LLM Use Cases
by: Cheng, Jingwen, et al.
Published: (2025)
by: Cheng, Jingwen, et al.
Published: (2025)
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
by: Pal, Soumyadeep, et al.
Published: (2025)
by: Pal, Soumyadeep, et al.
Published: (2025)
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
by: Xie, Linxi, et al.
Published: (2025)
by: Xie, Linxi, et al.
Published: (2025)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024)
by: Haimes, Jacob, et al.
Published: (2024)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
by: Cattan, Arie, et al.
Published: (2025)
by: Cattan, Arie, et al.
Published: (2025)
Similar Items
-
Personal Information Parroting in Language Models
by: Subramani, Nishant, et al.
Published: (2026) -
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
by: Ghate, Kshitish, et al.
Published: (2025) -
Two Types of AI Existential Risk: Decisive and Accumulative
by: Kasirzadeh, Atoosa
Published: (2024) -
Value Internalization: Learning and Generalizing from Social Reward
by: Rong, Frieda, et al.
Published: (2024) -
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
by: Muhamed, Aashiq, et al.
Published: (2024)