From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Fulay, Suyash, Zhu, Jocelyn, Bakker, Michiel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI and Collective Decisions: Strengthening Legitimacy and Losers' Consent
by: Fulay, Suyash, et al.
Published: (2026)
by: Fulay, Suyash, et al.
Published: (2026)
Agora: Teaching the Skill of Consensus-Finding with AI Personas Grounded in Human Voice
by: Ravi, Prerna, et al.
Published: (2026)
by: Ravi, Prerna, et al.
Published: (2026)
The Empty Chair: Using LLMs to Raise Missing Perspectives in Policy Deliberations
by: Fulay, Suyash, et al.
Published: (2025)
by: Fulay, Suyash, et al.
Published: (2025)
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
by: Yao, Xintong
Published: (2026)
by: Yao, Xintong
Published: (2026)
On the Relationship between Truth and Political Bias in Language Models
by: Fulay, Suyash, et al.
Published: (2024)
by: Fulay, Suyash, et al.
Published: (2024)
The Last Fingerprint: How Markdown Training Shapes LLM Prose
by: Freeburg, E. M.
Published: (2026)
by: Freeburg, E. M.
Published: (2026)
Implicit Bias in LLMs for Transgender Populations
by: Hirsch, Micaela, et al.
Published: (2026)
by: Hirsch, Micaela, et al.
Published: (2026)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
by: Vaccaro, Michelle, et al.
Published: (2026)
by: Vaccaro, Michelle, et al.
Published: (2026)
From Bias to Accountability: How the EU AI Act Confronts Challenges in European GeoAI Auditing
by: Matuszczyk, Natalia, et al.
Published: (2025)
by: Matuszczyk, Natalia, et al.
Published: (2025)
The Silent Curriculum: How Does LLM Monoculture Shape Educational Content and Its Accessibility?
by: Priyanshu, Aman, et al.
Published: (2024)
by: Priyanshu, Aman, et al.
Published: (2024)
Gender Bias of LLM in Economics: An Existentialism Perspective
by: Zhong, Hui, et al.
Published: (2024)
by: Zhong, Hui, et al.
Published: (2024)
AI Delegates with a Dual Focus: Ensuring Privacy and Strategic Self-Disclosure
by: Zhang, Zhiyang, et al.
Published: (2024)
by: Zhang, Zhiyang, et al.
Published: (2024)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
by: Nghiem, Huy, et al.
Published: (2026)
by: Nghiem, Huy, et al.
Published: (2026)
An Evaluation of Cultural Value Alignment in LLM
by: Sukiennik, Nicholas, et al.
Published: (2025)
by: Sukiennik, Nicholas, et al.
Published: (2025)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
by: Qu, Jinxian, et al.
Published: (2026)
by: Qu, Jinxian, et al.
Published: (2026)
From Bias Mitigation to Bias Negotiation: Governing Identity and Sociocultural Reasoning in Generative AI
by: Dunivin, Zackary Okun, et al.
Published: (2026)
by: Dunivin, Zackary Okun, et al.
Published: (2026)
Societal Alignment Frameworks Can Improve LLM Alignment
by: Stańczak, Karolina, et al.
Published: (2025)
by: Stańczak, Karolina, et al.
Published: (2025)
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
by: Wu, Addison J., et al.
Published: (2026)
by: Wu, Addison J., et al.
Published: (2026)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
by: Sun, Lihao, et al.
Published: (2025)
by: Sun, Lihao, et al.
Published: (2025)
Habermolt: Delegating Deliberation to AI Representatives
by: Low, Joseph, et al.
Published: (2026)
by: Low, Joseph, et al.
Published: (2026)
Path Dependence under Adaptive AI Delegation
by: Huang, Lingxiao, et al.
Published: (2026)
by: Huang, Lingxiao, et al.
Published: (2026)
Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews
by: Vasu, Sai Suresh Macharla, et al.
Published: (2025)
by: Vasu, Sai Suresh Macharla, et al.
Published: (2025)
How will advanced AI systems impact democracy?
by: Summerfield, Christopher, et al.
Published: (2024)
by: Summerfield, Christopher, et al.
Published: (2024)
Delegation and Verification Under AI
by: Huang, Lingxiao, et al.
Published: (2026)
by: Huang, Lingxiao, et al.
Published: (2026)
No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
by: Chand, Shireen, et al.
Published: (2025)
by: Chand, Shireen, et al.
Published: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
by: Jiang, Han, et al.
Published: (2025)
by: Jiang, Han, et al.
Published: (2025)
Diagnosing Korean-Language LLM Political Bias via Census-Grounded Agent Simulation
by: Kang, Sungwoo
Published: (2026)
by: Kang, Sungwoo
Published: (2026)
When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents
by: Wang, Zongwei, et al.
Published: (2026)
by: Wang, Zongwei, et al.
Published: (2026)
How Supply Chain Dependencies Complicate Bias Measurement and Accountability Attribution in AI Hiring Applications
by: Sharma, Gauri, et al.
Published: (2026)
by: Sharma, Gauri, et al.
Published: (2026)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
by: Sakhawat, Adib, et al.
Published: (2026)
by: Sakhawat, Adib, et al.
Published: (2026)
Moral Alignment for LLM Agents
by: Tennant, Elizaveta, et al.
Published: (2024)
by: Tennant, Elizaveta, et al.
Published: (2024)
Designing Algorithmic Delegates: The Role of Indistinguishability in Human-AI Handoff
by: Greenwood, Sophie, et al.
Published: (2025)
by: Greenwood, Sophie, et al.
Published: (2025)
Echoes of AI Harms: A Human-LLM Synergistic Framework for Bias-Driven Harm Anticipation
by: Tantalaki, Nicoleta, et al.
Published: (2025)
by: Tantalaki, Nicoleta, et al.
Published: (2025)
Visibility Allocation Systems: How Algorithmic Design Shapes Online Visibility and Societal Outcomes
by: Ionescu, Stefania, et al.
Published: (2025)
by: Ionescu, Stefania, et al.
Published: (2025)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
by: Brophy, Matthew
Published: (2025)
by: Brophy, Matthew
Published: (2025)
Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph
by: Tang, Zhenheng, et al.
Published: (2026)
by: Tang, Zhenheng, et al.
Published: (2026)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
by: Kwon, Jea, et al.
Published: (2025)
by: Kwon, Jea, et al.
Published: (2025)
Rethinking LLM Bias Probing Using Lessons from the Social Sciences
by: Morehouse, Kirsten N., et al.
Published: (2025)
by: Morehouse, Kirsten N., et al.
Published: (2025)
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation
by: Ghosh, Himel, et al.
Published: (2026)
by: Ghosh, Himel, et al.
Published: (2026)
Similar Items
-
AI and Collective Decisions: Strengthening Legitimacy and Losers' Consent
by: Fulay, Suyash, et al.
Published: (2026) -
Agora: Teaching the Skill of Consensus-Finding with AI Personas Grounded in Human Voice
by: Ravi, Prerna, et al.
Published: (2026) -
The Empty Chair: Using LLMs to Raise Missing Perspectives in Policy Deliberations
by: Fulay, Suyash, et al.
Published: (2025) -
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
by: Yao, Xintong
Published: (2026) -
On the Relationship between Truth and Political Bias in Language Models
by: Fulay, Suyash, et al.
Published: (2024)