LLM Safety Alignment is Divergence Estimation in Disguise

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Haldar, Rajdeep, Wang, Ziyi, Song, Qifan, Lin, Guang, Xing, Yue
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912661462056960
author Haldar, Rajdeep
Wang, Ziyi
Song, Qifan
Lin, Guang
Xing, Yue
author_facet Haldar, Rajdeep
Wang, Ziyi
Song, Qifan
Lin, Guang
Xing, Yue
contents We present a theoretical framework showing that popular LLM alignment methods, including RLHF and its variants, can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less preferred) distributions. This perspective explains the emergence of separation in the latent space between safe and harmful prompts after alignment. As an application of our general divergence framework, we propose KLDO, a novel KL divergence-based alignment method, and empirically validate its effectiveness. We further show that using compliance-refusal datasets, rather than standard preference-based datasets, leads to stronger separation and improved safety alignment. Finally, to quantify the separation effect, we propose a distance-based metric in the prompt representation space, which also acts as a statistically significant indicator for model safety.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00657
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Safety Alignment is Divergence Estimation in Disguise
Haldar, Rajdeep
Wang, Ziyi
Song, Qifan
Lin, Guang
Xing, Yue
Machine Learning
Artificial Intelligence
Computers and Society
We present a theoretical framework showing that popular LLM alignment methods, including RLHF and its variants, can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less preferred) distributions. This perspective explains the emergence of separation in the latent space between safe and harmful prompts after alignment. As an application of our general divergence framework, we propose KLDO, a novel KL divergence-based alignment method, and empirically validate its effectiveness. We further show that using compliance-refusal datasets, rather than standard preference-based datasets, leads to stronger separation and improved safety alignment. Finally, to quantify the separation effect, we propose a distance-based metric in the prompt representation space, which also acts as a statistically significant indicator for model safety.
title LLM Safety Alignment is Divergence Estimation in Disguise
topic Machine Learning
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2502.00657