Writer-R1: Enhancing Generative Writing in LLMs via Memory-augmented Replay Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Jihao, Zu, Shuaishuai, Ji, Zhiyuan, Zhou, Chunlai, Qin, Biao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915865956450304
author Zhao, Jihao
Zu, Shuaishuai
Ji, Zhiyuan
Zhou, Chunlai
Qin, Biao
author_facet Zhao, Jihao
Zu, Shuaishuai
Ji, Zhiyuan
Zhou, Chunlai
Qin, Biao
contents As a typical open-ended generation task, creative writing lacks verifiable reference answers, which has long constrained reward modeling and automatic evaluation due to high human annotation costs, evaluative bias, and coarse feedback signals. To address these challenges, this paper first designs a multi-agent collaborative workflow based on Grounded Theory, performing dimensional decomposition and hierarchical induction of the problem to dynamically produce interpretable and reusable fine-grained criteria. Furthermore, we propose the Memory-augmented Replay Policy Optimization (MRPO) algorithm: on the one hand, without additional training, MRPO guides models to engage in self-reflection based on dynamic criteria, enabling controlled iterative improvement; on the other hand, we adopt the training paradigm that combines supervised fine-tuning with reinforcement learning to convert evaluation criteria into reward signals, achieving end-to-end optimization. Experimental results demonstrate that the automatically constructed criteria achieve performance gains comparable to human annotations. Writer-R1-4B models trained with this approach outperform baselines across multiple creative writing tasks and surpass some 100B+ parameter open-source models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15061
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Writer-R1: Enhancing Generative Writing in LLMs via Memory-augmented Replay Policy Optimization
Zhao, Jihao
Zu, Shuaishuai
Ji, Zhiyuan
Zhou, Chunlai
Qin, Biao
Computation and Language
As a typical open-ended generation task, creative writing lacks verifiable reference answers, which has long constrained reward modeling and automatic evaluation due to high human annotation costs, evaluative bias, and coarse feedback signals. To address these challenges, this paper first designs a multi-agent collaborative workflow based on Grounded Theory, performing dimensional decomposition and hierarchical induction of the problem to dynamically produce interpretable and reusable fine-grained criteria. Furthermore, we propose the Memory-augmented Replay Policy Optimization (MRPO) algorithm: on the one hand, without additional training, MRPO guides models to engage in self-reflection based on dynamic criteria, enabling controlled iterative improvement; on the other hand, we adopt the training paradigm that combines supervised fine-tuning with reinforcement learning to convert evaluation criteria into reward signals, achieving end-to-end optimization. Experimental results demonstrate that the automatically constructed criteria achieve performance gains comparable to human annotations. Writer-R1-4B models trained with this approach outperform baselines across multiple creative writing tasks and surpass some 100B+ parameter open-source models.
title Writer-R1: Enhancing Generative Writing in LLMs via Memory-augmented Replay Policy Optimization
topic Computation and Language
url https://arxiv.org/abs/2603.15061