Saved in:
Bibliographic Details
Main Authors: Kiet, Huynh Trung, Minh, Dao Sy Duy, Nguyen, Tuan, Tran, Chi-Nguyen, Pham, Phu-Hoa, Quy, Nguyen Lam Phu, Han, The Anh, Tran-Thanh, Long
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.10843
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913139732250624
author Kiet, Huynh Trung
Minh, Dao Sy Duy
Nguyen, Tuan
Tran, Chi-Nguyen
Pham, Phu-Hoa
Quy, Nguyen Lam Phu
Han, The Anh
Tran-Thanh, Long
author_facet Kiet, Huynh Trung
Minh, Dao Sy Duy
Nguyen, Tuan
Tran, Chi-Nguyen
Pham, Phu-Hoa
Quy, Nguyen Lam Phu
Han, The Anh
Tran-Thanh, Long
contents Large language models increasingly mediate decisions that turn on moral judgement, yet a growing body of evidence shows that their implicit preferences are not culturally neutral. Existing cultural alignment methods either require per-country preference data and fine-tuning budgets or assume white-box access to model internals that commercial APIs do not expose. In this work, we focus on this realistic black-box, public-data-only regime and observe that within-country sociodemographic disagreement, not consensus, is the primary steering signal. We introduce DISCA (Disagreement-Informed Steering for Cultural Alignment), an inference-time method that instantiates each country as a panel of World-Values-Survey-grounded persona agents and converts their disagreement into a bounded, loss-averse logit correction. Across 20 countries and 7 open-weight backbones (2B--70B), DISCA reduces cultural misalignment on MultiTP by 10--24% on the six backbones >=3.8B, and 2--7% on open-ended scenarios, without changing any weights. Our results suggest that inference-time calibration is a scalable alternative to fine-tuning for serving the long tail of global moral preferences.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10843
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
Kiet, Huynh Trung
Minh, Dao Sy Duy
Nguyen, Tuan
Tran, Chi-Nguyen
Pham, Phu-Hoa
Quy, Nguyen Lam Phu
Han, The Anh
Tran-Thanh, Long
Computation and Language
Artificial Intelligence
Computers and Society
Large language models increasingly mediate decisions that turn on moral judgement, yet a growing body of evidence shows that their implicit preferences are not culturally neutral. Existing cultural alignment methods either require per-country preference data and fine-tuning budgets or assume white-box access to model internals that commercial APIs do not expose. In this work, we focus on this realistic black-box, public-data-only regime and observe that within-country sociodemographic disagreement, not consensus, is the primary steering signal. We introduce DISCA (Disagreement-Informed Steering for Cultural Alignment), an inference-time method that instantiates each country as a panel of World-Values-Survey-grounded persona agents and converts their disagreement into a bounded, loss-averse logit correction. Across 20 countries and 7 open-weight backbones (2B--70B), DISCA reduces cultural misalignment on MultiTP by 10--24% on the six backbones >=3.8B, and 2--7% on open-ended scenarios, without changing any weights. Our results suggest that inference-time calibration is a scalable alternative to fine-tuning for serving the long tail of global moral preferences.
title Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2605.10843