Misaligned by Reward: Socially Undesirable Preferences in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ghazaryan, Gayane, Dönmez, Esra |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
"They parted illusions -- they parted disclaim marinade": Misalignment as structural fidelity in LLMs
by: Costa, Mariana Lins
Published: (2025)
by: Costa, Mariana Lins
Published: (2025)
The Political Preferences of LLMs
by: Rozado, David
Published: (2024)
by: Rozado, David
Published: (2024)
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
by: Ghazaryan, Gayane, et al.
Published: (2024)
by: Ghazaryan, Gayane, et al.
Published: (2024)
neuralFOMO: Can LLMs Handle Being Second Best? Measuring Envy-Like Preferences in Multi-Agent Settings
by: Ramamoorthy, Arnav, et al.
Published: (2025)
by: Ramamoorthy, Arnav, et al.
Published: (2025)
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
by: Karr Jr., Jonathan A., et al.
Published: (2025)
by: Karr Jr., Jonathan A., et al.
Published: (2025)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
by: Wang, Peisong, et al.
Published: (2025)
by: Wang, Peisong, et al.
Published: (2025)
Temporal Preferences in Language Models for Long-Horizon Assistance
by: Mazyaki, Ali, et al.
Published: (2025)
by: Mazyaki, Ali, et al.
Published: (2025)
Measuring Political Preferences in AI Systems: An Integrative Approach
by: Rozado, David
Published: (2025)
by: Rozado, David
Published: (2025)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
by: Mavi, John, et al.
Published: (2024)
by: Mavi, John, et al.
Published: (2024)
The simulation of judgment in LLMs
by: Loru, Edoardo, et al.
Published: (2025)
by: Loru, Edoardo, et al.
Published: (2025)
Measuring Teaching with LLMs
by: Hardy, Michael
Published: (2025)
by: Hardy, Michael
Published: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
by: de Landa, Joseba Fernandez, et al.
Published: (2026)
by: de Landa, Joseba Fernandez, et al.
Published: (2026)
Growth First, Care Second? Tracing the Landscape of LLM Value Preferences in Everyday Dilemmas
by: Chen, Zhiyi, et al.
Published: (2026)
by: Chen, Zhiyi, et al.
Published: (2026)
Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers
by: Chakrabarty, Tuhin, et al.
Published: (2025)
by: Chakrabarty, Tuhin, et al.
Published: (2025)
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
by: Vida, Karina, et al.
Published: (2024)
by: Vida, Karina, et al.
Published: (2024)
Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
by: Ashkinaze, Joshua, et al.
Published: (2025)
by: Ashkinaze, Joshua, et al.
Published: (2025)
Moral Mazes in the Era of LLMs
by: Nguyen, Dang, et al.
Published: (2026)
by: Nguyen, Dang, et al.
Published: (2026)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
by: Wachter, Jasmin, et al.
Published: (2025)
by: Wachter, Jasmin, et al.
Published: (2025)
Interpretability Framework for LLMs in Undergraduate Calculus
by: Dakshit, Sagnik, et al.
Published: (2025)
by: Dakshit, Sagnik, et al.
Published: (2025)
On the Credibility of Evaluating LLMs using Survey Questions
by: Libovický, Jindřich
Published: (2026)
by: Libovický, Jindřich
Published: (2026)
Verbalizing LLMs' assumptions to explain and control sycophancy
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
by: Ghosh, Rajarshi, et al.
Published: (2025)
by: Ghosh, Rajarshi, et al.
Published: (2025)
ELEPHANT: Measuring and understanding social sycophancy in LLMs
by: Cheng, Myra, et al.
Published: (2025)
by: Cheng, Myra, et al.
Published: (2025)
BLIP: Facilitating the Exploration of Undesirable Consequences of Digital Technologies
by: Pang, Rock Yuren, et al.
Published: (2024)
by: Pang, Rock Yuren, et al.
Published: (2024)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
by: Cho, Jay Hyeon, et al.
Published: (2025)
by: Cho, Jay Hyeon, et al.
Published: (2025)
I Want to Break Free! Persuasion and Anti-Social Behavior of LLMs in Multi-Agent Settings with Social Hierarchy
by: Campedelli, Gian Maria, et al.
Published: (2024)
by: Campedelli, Gian Maria, et al.
Published: (2024)
Perceived Political Bias in LLMs Reduces Persuasive Abilities
by: DiGiuseppe, Matthew, et al.
Published: (2026)
by: DiGiuseppe, Matthew, et al.
Published: (2026)
Large Language Models (LLMs) as Agents for Augmented Democracy
by: Gudiño-Rosero, Jairo, et al.
Published: (2024)
by: Gudiño-Rosero, Jairo, et al.
Published: (2024)
Automated Assessment of Students' Code Comprehension using LLMs
by: Oli, Priti, et al.
Published: (2023)
by: Oli, Priti, et al.
Published: (2023)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
by: Allaham, Mowafak, et al.
Published: (2024)
by: Allaham, Mowafak, et al.
Published: (2024)
Explore the Potential of LLMs in Misinformation Detection: An Empirical Study
by: Chen, Mengyang, et al.
Published: (2023)
by: Chen, Mengyang, et al.
Published: (2023)
EtiCor++: Towards Understanding Etiquettical Bias in LLMs
by: Dwivedi, Ashutosh, et al.
Published: (2025)
by: Dwivedi, Ashutosh, et al.
Published: (2025)
EulerESG: Automating ESG Disclosure Analysis with LLMs
by: Ding, Yi, et al.
Published: (2025)
by: Ding, Yi, et al.
Published: (2025)
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
by: Dawson, Fiifi, et al.
Published: (2024)
by: Dawson, Fiifi, et al.
Published: (2024)
Towards Measuring and Modeling "Culture" in LLMs: A Survey
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human Feedback
by: Lerner, Emilia Agis, et al.
Published: (2024)
by: Lerner, Emilia Agis, et al.
Published: (2024)
Similar Items
-
"They parted illusions -- they parted disclaim marinade": Misalignment as structural fidelity in LLMs
by: Costa, Mariana Lins
Published: (2025) -
The Political Preferences of LLMs
by: Rozado, David
Published: (2024) -
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
by: Ghazaryan, Gayane, et al.
Published: (2024) -
neuralFOMO: Can LLMs Handle Being Second Best? Measuring Envy-Like Preferences in Multi-Agent Settings
by: Ramamoorthy, Arnav, et al.
Published: (2025) -
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
by: Karr Jr., Jonathan A., et al.
Published: (2025)