Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918338885582848 |
|---|---|
| author | Mukherjee, Subhojyoti Lai, Viet Dac Addanki, Raghavendra Rossi, Ryan Yoon, Seunghyun Bui, Trung Rao, Anup Subramanian, Jayakumar Kveton, Branislav |
| author_facet | Mukherjee, Subhojyoti Lai, Viet Dac Addanki, Raghavendra Rossi, Ryan Yoon, Seunghyun Bui, Trung Rao, Anup Subramanian, Jayakumar Kveton, Branislav |
| contents | Offline reinforcement learning (RL) is a variant of RL where the policy is learned from a previously collected dataset of trajectories and rewards. In our work, we propose a practical approach to offline RL with large language models (LLMs). We recast the problem as reward-weighted fine-tuning, which can be solved using similar techniques to supervised fine-tuning (SFT). To showcase the value of our approach, we apply it to learning short-horizon question-answering policies of a fixed length, where the agent reasons about potential answers or asks clarifying questions. Our work stands in a stark contrast to state-of-the-art methods in this domain, based on SFT and direct preference optimization, which have additional hyper-parameters and do not directly optimize for rewards. We compare to them empirically, and report major gains in both optimized rewards and language quality. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_06964 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization Mukherjee, Subhojyoti Lai, Viet Dac Addanki, Raghavendra Rossi, Ryan Yoon, Seunghyun Bui, Trung Rao, Anup Subramanian, Jayakumar Kveton, Branislav Computation and Language Machine Learning Offline reinforcement learning (RL) is a variant of RL where the policy is learned from a previously collected dataset of trajectories and rewards. In our work, we propose a practical approach to offline RL with large language models (LLMs). We recast the problem as reward-weighted fine-tuning, which can be solved using similar techniques to supervised fine-tuning (SFT). To showcase the value of our approach, we apply it to learning short-horizon question-answering policies of a fixed length, where the agent reasons about potential answers or asks clarifying questions. Our work stands in a stark contrast to state-of-the-art methods in this domain, based on SFT and direct preference optimization, which have additional hyper-parameters and do not directly optimize for rewards. We compare to them empirically, and report major gains in both optimized rewards and language quality. |
| title | Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2506.06964 |