From Coarse to Fine: Benchmarking and Reward Modeling for Writing-Centric Generation Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ren, Qingyu, Pan, Tianjun, Chen, Xingzhou, Wang, Xuhong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909003673501696
author Ren, Qingyu
Pan, Tianjun
Chen, Xingzhou
Wang, Xuhong
author_facet Ren, Qingyu
Pan, Tianjun
Chen, Xingzhou
Wang, Xuhong
contents Large language models have achieved remarkable progress in text generation but still struggle with generative writing tasks. In terms of evaluation, existing benchmarks evaluate writing reward models coarsely and fail to measure performance from the perspective of specific requirements. In terms of training, existing training methods either use LLM-as-a-judge approaches or train coarse-grained reward models, lacking fine-grained requirement-adherence reward modeling. To address these issues, we propose a fine-grained evaluation pipeline WEval for writing reward models and a fine-grained reinforcement learning training framework WRL. The evaluation data of WEval covers multiple task categories and requirement types, enabling systematic evaluation of writing reward models by measuring the correlation between the rankings of the reward model and gold rankings. WRL constructs positive and negative samples by selectively dropping instruction requirements, allowing for more precise reward model training. Experiments show that our models achieve substantial improvements across various writing benchmarks and exhibit strong generalization. The code and data are publicly available at \href{https://github.com/Rainier-rq1/From_Coarse_to_Fine}{https://github.com/Rainier-rq1/From\_Coarse\_to\_Fine}.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27453
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Coarse to Fine: Benchmarking and Reward Modeling for Writing-Centric Generation Tasks
Ren, Qingyu
Pan, Tianjun
Chen, Xingzhou
Wang, Xuhong
Computation and Language
Large language models have achieved remarkable progress in text generation but still struggle with generative writing tasks. In terms of evaluation, existing benchmarks evaluate writing reward models coarsely and fail to measure performance from the perspective of specific requirements. In terms of training, existing training methods either use LLM-as-a-judge approaches or train coarse-grained reward models, lacking fine-grained requirement-adherence reward modeling. To address these issues, we propose a fine-grained evaluation pipeline WEval for writing reward models and a fine-grained reinforcement learning training framework WRL. The evaluation data of WEval covers multiple task categories and requirement types, enabling systematic evaluation of writing reward models by measuring the correlation between the rankings of the reward model and gold rankings. WRL constructs positive and negative samples by selectively dropping instruction requirements, allowing for more precise reward model training. Experiments show that our models achieve substantial improvements across various writing benchmarks and exhibit strong generalization. The code and data are publicly available at \href{https://github.com/Rainier-rq1/From_Coarse_to_Fine}{https://github.com/Rainier-rq1/From\_Coarse\_to\_Fine}.
title From Coarse to Fine: Benchmarking and Reward Modeling for Writing-Centric Generation Tasks
topic Computation and Language
url https://arxiv.org/abs/2604.27453