TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Truong, Nguyen, Tien-Phat, Van, Linh Ngo, Nguyen, Duy Minh Ho, Doan, Khoa D., Le, Trung
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917495466622976
author Nguyen, Truong
Nguyen, Tien-Phat
Van, Linh Ngo
Nguyen, Duy Minh Ho
Doan, Khoa D.
Le, Trung
author_facet Nguyen, Truong
Nguyen, Tien-Phat
Van, Linh Ngo
Nguyen, Duy Minh Ho
Doan, Khoa D.
Le, Trung
contents Direct Preference Optimization (DPO) is a widely used RL-free method for aligning language models from pairwise preferences, but it models preferences over full sequences even though generation is driven by per-token decisions. Existing token-level extensions typically decompose a sequence-level Bradley-Terry objective across timesteps, leaving per-prefix (state-wise) optimality implicit. We study how to recover token-level preference optimality using only standard sequence-level pairwise comparisons. We introduce Token-level Bregman Preference Optimization (TBPO), which posits a token-level Bradley-Terry preference model over next-token actions conditioned on the prefix, and derive a Bregman-divergence density-ratio matching objective that generalizes the logistic/DPO loss while preserving the optimal policy induced by the token-level model and maintaining DPO-like simplicity. We introduce two instantiations: TBPO-Q, which explicitly learns a lightweight state baseline, and TBPO-A, which removes the baseline through advantage normalization. Across instruction following, helpfulness/harmlessness, and summarization benchmarks, TBPO improves alignment quality and training stability and increases output diversity relative to strong sequence-level and token-level baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12288
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
Nguyen, Truong
Nguyen, Tien-Phat
Van, Linh Ngo
Nguyen, Duy Minh Ho
Doan, Khoa D.
Le, Trung
Computation and Language
Artificial Intelligence
Direct Preference Optimization (DPO) is a widely used RL-free method for aligning language models from pairwise preferences, but it models preferences over full sequences even though generation is driven by per-token decisions. Existing token-level extensions typically decompose a sequence-level Bradley-Terry objective across timesteps, leaving per-prefix (state-wise) optimality implicit. We study how to recover token-level preference optimality using only standard sequence-level pairwise comparisons. We introduce Token-level Bregman Preference Optimization (TBPO), which posits a token-level Bradley-Terry preference model over next-token actions conditioned on the prefix, and derive a Bregman-divergence density-ratio matching objective that generalizes the logistic/DPO loss while preserving the optimal policy induced by the token-level model and maintaining DPO-like simplicity. We introduce two instantiations: TBPO-Q, which explicitly learns a lightweight state baseline, and TBPO-A, which removes the baseline through advantage normalization. Across instruction following, helpfulness/harmlessness, and summarization benchmarks, TBPO improves alignment quality and training stability and increases output diversity relative to strong sequence-level and token-level baselines.
title TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.12288