Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mukherjee, Subhojyoti, Lai, Viet Dac, Addanki, Raghavendra, Rossi, Ryan, Yoon, Seunghyun, Bui, Trung, Rao, Anup, Subramanian, Jayakumar, Kveton, Branislav
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918338885582848
author Mukherjee, Subhojyoti
Lai, Viet Dac
Addanki, Raghavendra
Rossi, Ryan
Yoon, Seunghyun
Bui, Trung
Rao, Anup
Subramanian, Jayakumar
Kveton, Branislav
author_facet Mukherjee, Subhojyoti
Lai, Viet Dac
Addanki, Raghavendra
Rossi, Ryan
Yoon, Seunghyun
Bui, Trung
Rao, Anup
Subramanian, Jayakumar
Kveton, Branislav
contents Offline reinforcement learning (RL) is a variant of RL where the policy is learned from a previously collected dataset of trajectories and rewards. In our work, we propose a practical approach to offline RL with large language models (LLMs). We recast the problem as reward-weighted fine-tuning, which can be solved using similar techniques to supervised fine-tuning (SFT). To showcase the value of our approach, we apply it to learning short-horizon question-answering policies of a fixed length, where the agent reasons about potential answers or asks clarifying questions. Our work stands in a stark contrast to state-of-the-art methods in this domain, based on SFT and direct preference optimization, which have additional hyper-parameters and do not directly optimize for rewards. We compare to them empirically, and report major gains in both optimized rewards and language quality.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06964
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
Mukherjee, Subhojyoti
Lai, Viet Dac
Addanki, Raghavendra
Rossi, Ryan
Yoon, Seunghyun
Bui, Trung
Rao, Anup
Subramanian, Jayakumar
Kveton, Branislav
Computation and Language
Machine Learning
Offline reinforcement learning (RL) is a variant of RL where the policy is learned from a previously collected dataset of trajectories and rewards. In our work, we propose a practical approach to offline RL with large language models (LLMs). We recast the problem as reward-weighted fine-tuning, which can be solved using similar techniques to supervised fine-tuning (SFT). To showcase the value of our approach, we apply it to learning short-horizon question-answering policies of a fixed length, where the agent reasons about potential answers or asks clarifying questions. Our work stands in a stark contrast to state-of-the-art methods in this domain, based on SFT and direct preference optimization, which have additional hyper-parameters and do not directly optimize for rewards. We compare to them empirically, and report major gains in both optimized rewards and language quality.
title Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.06964