Token-level Proximal Policy Optimization for Query Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ouyang, Yichen, Wang, Lu, Yang, Fangkai, Zhao, Pu, Huang, Chenghua, Liu, Jianfeng, Pang, Bochen, Yang, Yaming, Zhan, Yuefeng, Sun, Hao, Lin, Qingwei, Rajmohan, Saravan, Deng, Weiwei, Zhang, Dongmei, Sun, Feng, Zhang, Qi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929572073701376
author Ouyang, Yichen
Wang, Lu
Yang, Fangkai
Zhao, Pu
Huang, Chenghua
Liu, Jianfeng
Pang, Bochen
Yang, Yaming
Zhan, Yuefeng
Sun, Hao
Lin, Qingwei
Rajmohan, Saravan
Deng, Weiwei
Zhang, Dongmei
Sun, Feng
Zhang, Qi
author_facet Ouyang, Yichen
Wang, Lu
Yang, Fangkai
Zhao, Pu
Huang, Chenghua
Liu, Jianfeng
Pang, Bochen
Yang, Yaming
Zhan, Yuefeng
Sun, Hao
Lin, Qingwei
Rajmohan, Saravan
Deng, Weiwei
Zhang, Dongmei
Sun, Feng
Zhang, Qi
contents Query generation is a critical task for web search engines (e.g. Google, Bing) and recommendation systems. Recently, state-of-the-art query generation methods leverage Large Language Models (LLMs) for their strong capabilities in context understanding and text generation. However, they still face challenges in generating high-quality queries in terms of inferring user intent based on their web search interaction history. In this paper, we propose Token-level Proximal Policy Optimization (TPPO), a noval approach designed to empower LLMs perform better in query generation through fine-tuning. TPPO is based on the Reinforcement Learning from AI Feedback (RLAIF) paradigm, consisting of a token-level reward model and a token-level proximal policy optimization module to address the sparse reward challenge in traditional RLAIF frameworks. To evaluate the effectiveness and robustness of TPPO, we conducted experiments on both open-source dataset and an industrial dataset that was collected from a globally-used search engine. The experimental results demonstrate that TPPO significantly improves the performance of query generation for LLMs and outperforms its existing competitors.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00722
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Token-level Proximal Policy Optimization for Query Generation
Ouyang, Yichen
Wang, Lu
Yang, Fangkai
Zhao, Pu
Huang, Chenghua
Liu, Jianfeng
Pang, Bochen
Yang, Yaming
Zhan, Yuefeng
Sun, Hao
Lin, Qingwei
Rajmohan, Saravan
Deng, Weiwei
Zhang, Dongmei
Sun, Feng
Zhang, Qi
Machine Learning
Query generation is a critical task for web search engines (e.g. Google, Bing) and recommendation systems. Recently, state-of-the-art query generation methods leverage Large Language Models (LLMs) for their strong capabilities in context understanding and text generation. However, they still face challenges in generating high-quality queries in terms of inferring user intent based on their web search interaction history. In this paper, we propose Token-level Proximal Policy Optimization (TPPO), a noval approach designed to empower LLMs perform better in query generation through fine-tuning. TPPO is based on the Reinforcement Learning from AI Feedback (RLAIF) paradigm, consisting of a token-level reward model and a token-level proximal policy optimization module to address the sparse reward challenge in traditional RLAIF frameworks. To evaluate the effectiveness and robustness of TPPO, we conducted experiments on both open-source dataset and an industrial dataset that was collected from a globally-used search engine. The experimental results demonstrate that TPPO significantly improves the performance of query generation for LLMs and outperforms its existing competitors.
title Token-level Proximal Policy Optimization for Query Generation
topic Machine Learning
url https://arxiv.org/abs/2411.00722