From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Zhihan, Wu, Jiele, Cui, Wenqian, Zhang, Yifei, Hu, Minda, Wang, Yufei, King, Irwin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918064231022592
author Guo, Zhihan
Wu, Jiele
Cui, Wenqian
Zhang, Yifei
Hu, Minda
Wang, Yufei
King, Irwin
author_facet Guo, Zhihan
Wu, Jiele
Cui, Wenqian
Zhang, Yifei
Hu, Minda
Wang, Yufei
King, Irwin
contents Current research on long-form context in Large Language Models (LLMs) primarily focuses on the understanding of long-contexts, the Open-ended Long Text Generation (Open-LTG) remains insufficiently explored. Training a long-context generation model requires curation of gold standard reference data, which is typically nonexistent for informative Open-LTG tasks. However, previous methods only utilize general assessments as reward signals, which limits accuracy. To bridge this gap, we introduce ProxyReward, an innovative reinforcement learning (RL) based framework, which includes a dataset and a reward signal computation method. Firstly, ProxyReward Dataset generation is accomplished through simple prompts that enables the model to create automatically, obviating extensive labeled data or significant manual effort. Secondly, ProxyReward Signal offers a targeted evaluation of information comprehensiveness and accuracy for specific questions. The experimental results indicate that our method ProxyReward surpasses even GPT-4-Turbo. It can significantly enhance performance by 20% on the Open-LTG task when training widely used open-source models, while also surpassing the LLM-as-a-Judge approach. Our work presents effective methods to enhance the ability of LLMs to address complex open-ended questions posed by human.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation
Guo, Zhihan
Wu, Jiele
Cui, Wenqian
Zhang, Yifei
Hu, Minda
Wang, Yufei
King, Irwin
Computation and Language
Artificial Intelligence
Current research on long-form context in Large Language Models (LLMs) primarily focuses on the understanding of long-contexts, the Open-ended Long Text Generation (Open-LTG) remains insufficiently explored. Training a long-context generation model requires curation of gold standard reference data, which is typically nonexistent for informative Open-LTG tasks. However, previous methods only utilize general assessments as reward signals, which limits accuracy. To bridge this gap, we introduce ProxyReward, an innovative reinforcement learning (RL) based framework, which includes a dataset and a reward signal computation method. Firstly, ProxyReward Dataset generation is accomplished through simple prompts that enables the model to create automatically, obviating extensive labeled data or significant manual effort. Secondly, ProxyReward Signal offers a targeted evaluation of information comprehensiveness and accuracy for specific questions. The experimental results indicate that our method ProxyReward surpasses even GPT-4-Turbo. It can significantly enhance performance by 20% on the Open-LTG task when training widely used open-source models, while also surpassing the LLM-as-a-Judge approach. Our work presents effective methods to enhance the ability of LLMs to address complex open-ended questions posed by human.
title From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.16024