Distributionally Robust Token Optimization in RLHF

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jin, Yeping, Hu, Jiaming, Paschalidis, Ioannis Ch.
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918494409326592
author Jin, Yeping
Hu, Jiaming
Paschalidis, Ioannis Ch.
author_facet Jin, Yeping
Hu, Jiaming
Paschalidis, Ioannis Ch.
contents Large Language Models (LLMs) tend to respond correctly to prompts that align well with the data they were trained and fine-tuned on. Yet, small shifts in wording, format, or language can trigger surprisingly large failures, especially on multi-step reasoning problems. To address this problem, we propose a Distributionally Robust Token Optimization (DRTO) approach, which combines token-level Reinforcement Learning from Human Feedback (RLHF) with Distributionally Robust Optimization (DRO). DRTO constructs f-divergence ambiguity sets over span-level actor losses, providing a principled way to emphasize difficult response segments during policy optimization. Empirically, DRTO enhances consistency under distribution shifts in multiple reasoning benchmarks among different tasks, achieving $+4.4$ percentage points on MATH-500 and $+2.7$ percentage points on LiveCodeBench over standard RTO.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08577
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Distributionally Robust Token Optimization in RLHF
Jin, Yeping
Hu, Jiaming
Paschalidis, Ioannis Ch.
Machine Learning
Artificial Intelligence
Large Language Models (LLMs) tend to respond correctly to prompts that align well with the data they were trained and fine-tuned on. Yet, small shifts in wording, format, or language can trigger surprisingly large failures, especially on multi-step reasoning problems. To address this problem, we propose a Distributionally Robust Token Optimization (DRTO) approach, which combines token-level Reinforcement Learning from Human Feedback (RLHF) with Distributionally Robust Optimization (DRO). DRTO constructs f-divergence ambiguity sets over span-level actor losses, providing a principled way to emphasize difficult response segments during policy optimization. Empirically, DRTO enhances consistency under distribution shifts in multiple reasoning benchmarks among different tasks, achieving $+4.4$ percentage points on MATH-500 and $+2.7$ percentage points on LiveCodeBench over standard RTO.
title Distributionally Robust Token Optimization in RLHF
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.08577