TriAlign: Towards Universal Truth Consistency in Personalized LLM Alignment

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nguyen, Thi-Nhung, Luo, Linhao, Omari, Rollin, Kim, Junae, Vu, Thuy-Trang, Phung, Dinh
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913178106986496
author Nguyen, Thi-Nhung
Luo, Linhao
Omari, Rollin
Kim, Junae
Vu, Thuy-Trang
Phung, Dinh
author_facet Nguyen, Thi-Nhung
Luo, Linhao
Omari, Rollin
Kim, Junae
Vu, Thuy-Trang
Phung, Dinh
contents Personalized large language models adapt responses to users' preferences and social attributes, but can introduce substantial universal truth inconsistencies across social groups, where some groups systematically receive less accurate responses on objective tasks. Existing alignment methods either ignore personalization or mainly focus on subjective preference alignment, largely overlooking fairness and consistency in universal truths. To address this gap, we study Truth-Invariant Alignment (TIA), an alignment problem for personalized LLMs that aims to ensure universal truths remain consistent across social groups while preserving personalization. We propose TriAlign, the first offline multi-agent reinforcement learning (MARL) framework for TIA, where each social group is modeled as an agent interacting. TriAlign jointly optimizes universal truth accuracy, cross-group truth consistency, and personalization through a fairness-aware objective and an explicit inconsistency penalty. Experiments across diverse benchmarks demonstrate that TriAlign achieves a stronger balance among these three objectives than strong baselines, reducing universal truth disparities across social groups while improving both objective task performance and personalization quality.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01755
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TriAlign: Towards Universal Truth Consistency in Personalized LLM Alignment
Nguyen, Thi-Nhung
Luo, Linhao
Omari, Rollin
Kim, Junae
Vu, Thuy-Trang
Phung, Dinh
Artificial Intelligence
Computation and Language
Personalized large language models adapt responses to users' preferences and social attributes, but can introduce substantial universal truth inconsistencies across social groups, where some groups systematically receive less accurate responses on objective tasks. Existing alignment methods either ignore personalization or mainly focus on subjective preference alignment, largely overlooking fairness and consistency in universal truths. To address this gap, we study Truth-Invariant Alignment (TIA), an alignment problem for personalized LLMs that aims to ensure universal truths remain consistent across social groups while preserving personalization. We propose TriAlign, the first offline multi-agent reinforcement learning (MARL) framework for TIA, where each social group is modeled as an agent interacting. TriAlign jointly optimizes universal truth accuracy, cross-group truth consistency, and personalization through a fairness-aware objective and an explicit inconsistency penalty. Experiments across diverse benchmarks demonstrate that TriAlign achieves a stronger balance among these three objectives than strong baselines, reducing universal truth disparities across social groups while improving both objective task performance and personalization quality.
title TriAlign: Towards Universal Truth Consistency in Personalized LLM Alignment
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2606.01755