Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xinnan, Li, Chenliang, Zeng, Siliang, Li, Jiaxiang, Wang, Zhongruo, Lin, Kaixiang, Lu, Songtao, Garcia, Alfredo, Hong, Mingyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913923455778816
author Zhang, Xinnan
Li, Chenliang
Zeng, Siliang
Li, Jiaxiang
Wang, Zhongruo
Lin, Kaixiang
Lu, Songtao
Garcia, Alfredo
Hong, Mingyi
author_facet Zhang, Xinnan
Li, Chenliang
Zeng, Siliang
Li, Jiaxiang
Wang, Zhongruo
Lin, Kaixiang
Lu, Songtao
Garcia, Alfredo
Hong, Mingyi
contents Aligning large language models (LLMs) with human preferences usually requires fine-tuning methods such as RLHF and DPO. These methods directly optimize the model parameters, so they cannot be used in test-time to improve model performance, nor are they applicable when the model weights are not accessible. In contrast, test-time methods sidestep weight updates by leveraging reward functions to guide and improve output quality. However, they incur high inference costs, and their one-shot guidance is often based on imperfect reward or value functions, leading to suboptimal outputs. In this work, we present a method named Iterative Reweight-then-Optimize (IRO), a reinforcement learning (RL) framework that performs RL-style alignment of the (frozen) base model without touching its parameters. During training, each iteration (i) samples candidates from the base model, (ii) resamples using current value functions, and (iii) trains a new lightweight value function that guides the next decoding pass. At test time, the value functions are used to guide the base model generation via a search-based optimization process. Notably, users can apply IRO to align a model on their own dataset, similar to OpenAI's reinforcement fine-tuning (RFT), but without requiring access to the model weights.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17828
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach
Zhang, Xinnan
Li, Chenliang
Zeng, Siliang
Li, Jiaxiang
Wang, Zhongruo
Lin, Kaixiang
Lu, Songtao
Garcia, Alfredo
Hong, Mingyi
Machine Learning
Artificial Intelligence
Computation and Language
Aligning large language models (LLMs) with human preferences usually requires fine-tuning methods such as RLHF and DPO. These methods directly optimize the model parameters, so they cannot be used in test-time to improve model performance, nor are they applicable when the model weights are not accessible. In contrast, test-time methods sidestep weight updates by leveraging reward functions to guide and improve output quality. However, they incur high inference costs, and their one-shot guidance is often based on imperfect reward or value functions, leading to suboptimal outputs. In this work, we present a method named Iterative Reweight-then-Optimize (IRO), a reinforcement learning (RL) framework that performs RL-style alignment of the (frozen) base model without touching its parameters. During training, each iteration (i) samples candidates from the base model, (ii) resamples using current value functions, and (iii) trains a new lightweight value function that guides the next decoding pass. At test time, the value functions are used to guide the base model generation via a search-based optimization process. Notably, users can apply IRO to align a model on their own dataset, similar to OpenAI's reinforcement fine-tuning (RFT), but without requiring access to the model weights.
title Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.17828