DocReward: A Document Reward Model for Structuring and Stylizing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Junpeng, Zhao, Yuzhong, Cao, Bowen, Ding, Jiayu, Jia, Yilin, Lv, Tengchao, Huang, Yupan, Wu, Wenshan, Huang, Shaohan, Yang, Nan, Dong, Li, Cui, Lei, Ge, Tao, Wang, Xun, Jiao, Huitian, Mao, Sun, Kartik, FNU, Chen, Si-Qing, Lam, Wai, Wei, Furu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913138878709760
author Liu, Junpeng
Zhao, Yuzhong
Cao, Bowen
Ding, Jiayu
Jia, Yilin
Lv, Tengchao
Huang, Yupan
Wu, Wenshan
Huang, Shaohan
Yang, Nan
Dong, Li
Cui, Lei
Ge, Tao
Wang, Xun
Jiao, Huitian
Mao, Sun
Kartik, FNU
Chen, Si-Qing
Lam, Wai
Wei, Furu
author_facet Liu, Junpeng
Zhao, Yuzhong
Cao, Bowen
Ding, Jiayu
Jia, Yilin
Lv, Tengchao
Huang, Yupan
Wu, Wenshan
Huang, Shaohan
Yang, Nan
Dong, Li
Cui, Lei
Ge, Tao
Wang, Xun
Jiao, Huitian
Mao, Sun
Kartik, FNU
Chen, Si-Qing
Lam, Wai
Wei, Furu
contents Recent agentic workflows automate professional document generation but focus narrowly on textual quality, overlooking structural and stylistic professionalism, which is equally critical for readability. This gap stems mainly from a lack of effective reward models capable of guiding agents toward producing documents with high structural and stylistic professionalism. We introduce DocReward, a document reward model that evaluates documents based on their structure and style. To achieve this, we propose a textual-quality-agnostic framework that ensures assessments are not confounded by content quality, and construct DocPair, a dataset of 117K paired documents covering 32 domains and 267 types. Each pair shares identical content but differs in structural and stylistic professionalism. DocReward is trained using the Bradley-Terry loss. On a manually annotated benchmark, DocReward outperforms GPT-5 by 14.6 percentage points in the same setting. Reinforcement learning experiments further show that DocReward effectively guides agents toward generating documents with consistently higher structural and stylistic professionalism, highlighting its practical utility.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DocReward: A Document Reward Model for Structuring and Stylizing
Liu, Junpeng
Zhao, Yuzhong
Cao, Bowen
Ding, Jiayu
Jia, Yilin
Lv, Tengchao
Huang, Yupan
Wu, Wenshan
Huang, Shaohan
Yang, Nan
Dong, Li
Cui, Lei
Ge, Tao
Wang, Xun
Jiao, Huitian
Mao, Sun
Kartik, FNU
Chen, Si-Qing
Lam, Wai
Wei, Furu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Recent agentic workflows automate professional document generation but focus narrowly on textual quality, overlooking structural and stylistic professionalism, which is equally critical for readability. This gap stems mainly from a lack of effective reward models capable of guiding agents toward producing documents with high structural and stylistic professionalism. We introduce DocReward, a document reward model that evaluates documents based on their structure and style. To achieve this, we propose a textual-quality-agnostic framework that ensures assessments are not confounded by content quality, and construct DocPair, a dataset of 117K paired documents covering 32 domains and 267 types. Each pair shares identical content but differs in structural and stylistic professionalism. DocReward is trained using the Bradley-Terry loss. On a manually annotated benchmark, DocReward outperforms GPT-5 by 14.6 percentage points in the same setting. Reinforcement learning experiments further show that DocReward effectively guides agents toward generating documents with consistently higher structural and stylistic professionalism, highlighting its practical utility.
title DocReward: A Document Reward Model for Structuring and Stylizing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.11391