InSPO: Unlocking Intrinsic Self-Reflection for LLM Preference Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Yu, Lan, Tian, Qi, Zhengling
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908822016098304
author Li, Yu
Lan, Tian
Qi, Zhengling
author_facet Li, Yu
Lan, Tian
Qi, Zhengling
contents Direct Preference Optimization (DPO) and its variants have become standard for aligning Large Language Models due to their simplicity and offline stability. However, we identify two fundamental limitations. First, the optimal policy depends on arbitrary modeling choices (scalarization function, reference policy), yielding behavior reflecting parameterization artifacts rather than true preferences. Second, treating response generation in isolation fails to leverage comparative information in pairwise data, leaving the model's capacity for intrinsic self-reflection untapped. To address it, we propose Intrinsic Self-reflective Preference Optimization (InSPO), deriving a globally optimal policy conditioning on both context and alternative responses. We prove this formulation superior to DPO/RLHF while guaranteeing invariance to scalarization and reference choices. InSPO serves as a plug-and-play enhancement without architectural changes or inference overhead. Experiments demonstrate consistent improvements in win rates and length-controlled metrics, validating that unlocking self-reflection yields more robust, human-aligned LLMs. Our Code is available at https://github.com/Skylanding/InSPO.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23126
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InSPO: Unlocking Intrinsic Self-Reflection for LLM Preference Optimization
Li, Yu
Lan, Tian
Qi, Zhengling
Artificial Intelligence
Machine Learning
Direct Preference Optimization (DPO) and its variants have become standard for aligning Large Language Models due to their simplicity and offline stability. However, we identify two fundamental limitations. First, the optimal policy depends on arbitrary modeling choices (scalarization function, reference policy), yielding behavior reflecting parameterization artifacts rather than true preferences. Second, treating response generation in isolation fails to leverage comparative information in pairwise data, leaving the model's capacity for intrinsic self-reflection untapped. To address it, we propose Intrinsic Self-reflective Preference Optimization (InSPO), deriving a globally optimal policy conditioning on both context and alternative responses. We prove this formulation superior to DPO/RLHF while guaranteeing invariance to scalarization and reference choices. InSPO serves as a plug-and-play enhancement without architectural changes or inference overhead. Experiments demonstrate consistent improvements in win rates and length-controlled metrics, validating that unlocking self-reflection yields more robust, human-aligned LLMs. Our Code is available at https://github.com/Skylanding/InSPO.
title InSPO: Unlocking Intrinsic Self-Reflection for LLM Preference Optimization
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.23126