HAL: Inducing Human-likeness in LLMs with Alignment

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hasan, Masum, Zhao, Junjie, Hoque, Ehsan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911357480206336
author Hasan, Masum
Zhao, Junjie
Hoque, Ehsan
author_facet Hasan, Masum
Zhao, Junjie
Hoque, Ehsan
contents Conversational human-likeness plays a central role in human-AI interaction, yet it has remained difficult to define, measure, and optimize. As a result, improvements in human-like behavior are largely driven by scale or broad supervised training, rather than targeted alignment. We introduce Human Aligning LLMs (HAL), a framework for aligning language models to conversational human-likeness using an interpretable, data-driven reward. HAL derives explicit conversational traits from contrastive dialogue data, combines them into a compact scalar score, and uses this score as a transparent reward signal for alignment with standard preference optimization methods. Using this approach, we align models of varying sizes without affecting their overall performance. In large-scale human evaluations, models aligned with HAL are more frequently perceived as human-like in conversation. Because HAL operates over explicit, interpretable traits, it enables inspection of alignment behavior and diagnosis of unintended effects. More broadly, HAL demonstrates how soft, qualitative properties of language--previously outside the scope for alignment--can be made measurable and aligned in an interpretable and explainable way.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02813
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HAL: Inducing Human-likeness in LLMs with Alignment
Hasan, Masum
Zhao, Junjie
Hoque, Ehsan
Artificial Intelligence
Computation and Language
Machine Learning
Conversational human-likeness plays a central role in human-AI interaction, yet it has remained difficult to define, measure, and optimize. As a result, improvements in human-like behavior are largely driven by scale or broad supervised training, rather than targeted alignment. We introduce Human Aligning LLMs (HAL), a framework for aligning language models to conversational human-likeness using an interpretable, data-driven reward. HAL derives explicit conversational traits from contrastive dialogue data, combines them into a compact scalar score, and uses this score as a transparent reward signal for alignment with standard preference optimization methods. Using this approach, we align models of varying sizes without affecting their overall performance. In large-scale human evaluations, models aligned with HAL are more frequently perceived as human-like in conversation. Because HAL operates over explicit, interpretable traits, it enables inspection of alignment behavior and diagnosis of unintended effects. More broadly, HAL demonstrates how soft, qualitative properties of language--previously outside the scope for alignment--can be made measurable and aligned in an interpretable and explainable way.
title HAL: Inducing Human-likeness in LLMs with Alignment
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.02813