Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zhipeng, Zhou, Kun, Zhao, Wayne Xin, Wan, Junchen, Zhang, Fuzheng, Zhang, Di, Wen, Ji-Rong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916289133412352
author Chen, Zhipeng
Zhou, Kun
Zhao, Wayne Xin
Wan, Junchen
Zhang, Fuzheng
Zhang, Di
Wen, Ji-Rong
author_facet Chen, Zhipeng
Zhou, Kun
Zhao, Wayne Xin
Wan, Junchen
Zhang, Fuzheng
Zhang, Di
Wen, Ji-Rong
contents Reinforcement learning (RL) has been widely used in training large language models (LLMs) for preventing unexpected outputs, eg reducing harmfulness and errors. However, existing RL methods mostly adopt the instance-level reward, which is unable to provide fine-grained supervision for complex reasoning tasks, and can not focus on the few key tokens that lead to the incorrectness. To address it, we propose a new RL method named RLMEC that incorporates a generative model as the reward model, which is trained by the erroneous solution rewriting task under the minimum editing constraint, and can produce token-level rewards for RL training. Based on the generative reward model, we design the token-level RL objective for training and an imitation-based regularization for stabilizing RL process. And the both objectives focus on the learning of the key tokens for the erroneous solution, reducing the effect of other unimportant tokens. The experiment results on mathematical tasks and question-answering tasks have demonstrated the effectiveness of our approach. Our code and data are available at https://github.com/RUCAIBox/RLMEC.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06081
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint
Chen, Zhipeng
Zhou, Kun
Zhao, Wayne Xin
Wan, Junchen
Zhang, Fuzheng
Zhang, Di
Wen, Ji-Rong
Computation and Language
Reinforcement learning (RL) has been widely used in training large language models (LLMs) for preventing unexpected outputs, eg reducing harmfulness and errors. However, existing RL methods mostly adopt the instance-level reward, which is unable to provide fine-grained supervision for complex reasoning tasks, and can not focus on the few key tokens that lead to the incorrectness. To address it, we propose a new RL method named RLMEC that incorporates a generative model as the reward model, which is trained by the erroneous solution rewriting task under the minimum editing constraint, and can produce token-level rewards for RL training. Based on the generative reward model, we design the token-level RL objective for training and an imitation-based regularization for stabilizing RL process. And the both objectives focus on the learning of the key tokens for the erroneous solution, reducing the effect of other unimportant tokens. The experiment results on mathematical tasks and question-answering tasks have demonstrated the effectiveness of our approach. Our code and data are available at https://github.com/RUCAIBox/RLMEC.
title Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint
topic Computation and Language
url https://arxiv.org/abs/2401.06081