Reward Models Inherit Value Biases from Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Christian, Brian, Thompson, Jessica A. F., Yang, Elle Michelle, Adam, Vincent, Kirk, Hannah Rose, Summerfield, Christopher, Dumbalska, Tsvetomira |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reward Model Interpretability via Optimal and Pessimal Tokens
von: Christian, Brian, et al.
Veröffentlicht: (2025)
von: Christian, Brian, et al.
Veröffentlicht: (2025)
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
von: Khandelwal, Khyati, et al.
Veröffentlicht: (2023)
von: Khandelwal, Khyati, et al.
Veröffentlicht: (2023)
Measuring and Mitigating Persona Distortions from AI Writing Assistance
von: Röttger, Paul, et al.
Veröffentlicht: (2026)
von: Röttger, Paul, et al.
Veröffentlicht: (2026)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
von: Elle
Veröffentlicht: (2025)
von: Elle
Veröffentlicht: (2025)
Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
von: Christian, Brian, et al.
Veröffentlicht: (2026)
von: Christian, Brian, et al.
Veröffentlicht: (2026)
Zero-shot counting with a dual-stream neural network model
von: Thompson, Jessica A. F., et al.
Veröffentlicht: (2024)
von: Thompson, Jessica A. F., et al.
Veröffentlicht: (2024)
The Life Cycle of Large Language Models: A Review of Biases in Education
von: Lee, Jinsook, et al.
Veröffentlicht: (2024)
von: Lee, Jinsook, et al.
Veröffentlicht: (2024)
DoDo Learning: DOmain-DemOgraphic Transfer in Language Models for Detecting Abuse Targeted at Public Figures
von: Williams, Angus R., et al.
Veröffentlicht: (2023)
von: Williams, Angus R., et al.
Veröffentlicht: (2023)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2026)
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2026)
The Biased Samaritan: LLM biases in Perceived Kindness
von: Fagan, Jack H, et al.
Veröffentlicht: (2025)
von: Fagan, Jack H, et al.
Veröffentlicht: (2025)
Quantifying Gender Biases Towards Politicians on Reddit
von: Marjanovic, Sara, et al.
Veröffentlicht: (2021)
von: Marjanovic, Sara, et al.
Veröffentlicht: (2021)
Generative Language Models Exhibit Social Identity Biases
von: Hu, Tiancheng, et al.
Veröffentlicht: (2023)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2023)
The Silicon Ceiling: Auditing GPT's Race and Gender Biases in Hiring
von: Armstrong, Lena, et al.
Veröffentlicht: (2024)
von: Armstrong, Lena, et al.
Veröffentlicht: (2024)
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
von: Weissburg, Iain, et al.
Veröffentlicht: (2024)
von: Weissburg, Iain, et al.
Veröffentlicht: (2024)
Revealing and Reducing Gender Biases in Vision and Language Assistants (VLAs)
von: Girrbach, Leander, et al.
Veröffentlicht: (2024)
von: Girrbach, Leander, et al.
Veröffentlicht: (2024)
Beyond the Binary: Capturing Diverse Preferences With Reward Regularization
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2024)
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2024)
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
von: Gautam, Vagrant, et al.
Veröffentlicht: (2024)
von: Gautam, Vagrant, et al.
Veröffentlicht: (2024)
The Algorithmic Unconscious: Structural Mechanisms and Implicit Biases in Large Language Models
von: Boisnard, Philippe
Veröffentlicht: (2026)
von: Boisnard, Philippe
Veröffentlicht: (2026)
Toward Automated Detection of Biased Social Signals from the Content of Clinical Conversations
von: Chen, Feng, et al.
Veröffentlicht: (2024)
von: Chen, Feng, et al.
Veröffentlicht: (2024)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2025)
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2025)
Rejected Dialects: Biases Against African American Language in Reward Models
von: Mire, Joel, et al.
Veröffentlicht: (2025)
von: Mire, Joel, et al.
Veröffentlicht: (2025)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
von: Ghosh, Rajarshi, et al.
Veröffentlicht: (2025)
von: Ghosh, Rajarshi, et al.
Veröffentlicht: (2025)
Reducing Biases towards Minoritized Populations in Medical Curricular Content via Artificial Intelligence for Fairer Health Outcomes
von: Salavati, Chiman, et al.
Veröffentlicht: (2024)
von: Salavati, Chiman, et al.
Veröffentlicht: (2024)
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
von: Karr Jr., Jonathan A., et al.
Veröffentlicht: (2025)
von: Karr Jr., Jonathan A., et al.
Veröffentlicht: (2025)
A Systematic Analysis of Biases in Large Language Models
von: Zhang, Xulang, et al.
Veröffentlicht: (2025)
von: Zhang, Xulang, et al.
Veröffentlicht: (2025)
Understanding Intrinsic Socioeconomic Biases in Large Language Models
von: Arzaghi, Mina, et al.
Veröffentlicht: (2024)
von: Arzaghi, Mina, et al.
Veröffentlicht: (2024)
The Balancing Act: Unmasking and Alleviating ASR Biases in Portuguese
von: Kulkarni, Ajinkya, et al.
Veröffentlicht: (2024)
von: Kulkarni, Ajinkya, et al.
Veröffentlicht: (2024)
A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models
von: Pfohl, Stephen R., et al.
Veröffentlicht: (2024)
von: Pfohl, Stephen R., et al.
Veröffentlicht: (2024)
WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models
von: Zhao, Wenlong, et al.
Veröffentlicht: (2024)
von: Zhao, Wenlong, et al.
Veröffentlicht: (2024)
PRISM: A Methodology for Auditing Biases in Large Language Models
von: Azzopardi, Leif, et al.
Veröffentlicht: (2024)
von: Azzopardi, Leif, et al.
Veröffentlicht: (2024)
Defining bias in AI-systems: Biased models are fair models
von: Lindloff, Chiara, et al.
Veröffentlicht: (2025)
von: Lindloff, Chiara, et al.
Veröffentlicht: (2025)
Exploring Gender Biases in Language Patterns of Human-Conversational Agent Conversations
von: Liu, Weizi
Veröffentlicht: (2024)
von: Liu, Weizi
Veröffentlicht: (2024)
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
von: Jahara, Fatima, et al.
Veröffentlicht: (2025)
von: Jahara, Fatima, et al.
Veröffentlicht: (2025)
The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas
von: Wu, Ya, et al.
Veröffentlicht: (2025)
von: Wu, Ya, et al.
Veröffentlicht: (2025)
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
von: Wu, Addison J., et al.
Veröffentlicht: (2025)
von: Wu, Addison J., et al.
Veröffentlicht: (2025)
Improving Cross-Cultural Survey Simulation with Calibrated Value Personas
von: Abels, Axel, et al.
Veröffentlicht: (2026)
von: Abels, Axel, et al.
Veröffentlicht: (2026)
Cultural Value Differences of LLMs: Prompt, Language, and Model Size
von: Zhong, Qishuai, et al.
Veröffentlicht: (2024)
von: Zhong, Qishuai, et al.
Veröffentlicht: (2024)
MoVa: Towards Generalizable Classification of Human Morals and Values
von: Chen, Ziyu, et al.
Veröffentlicht: (2025)
von: Chen, Ziyu, et al.
Veröffentlicht: (2025)
Language Agents as Digital Representatives in Collective Decision-Making
von: Jarrett, Daniel, et al.
Veröffentlicht: (2025)
von: Jarrett, Daniel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reward Model Interpretability via Optimal and Pessimal Tokens
von: Christian, Brian, et al.
Veröffentlicht: (2025) -
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
von: Khandelwal, Khyati, et al.
Veröffentlicht: (2023) -
Measuring and Mitigating Persona Distortions from AI Writing Assistance
von: Röttger, Paul, et al.
Veröffentlicht: (2026) -
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025) -
Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
von: Elle
Veröffentlicht: (2025)