HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Andrew Zhuoer, Wang, Cunxiang, Luo, Yu, Fan, Lin, Zhou, Yilin, Wang, Zikang, Gu, Xiaotao, Tang, Jie, Wang, Hongning, Huang, Minlie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908982801596416
author Feng, Andrew Zhuoer
Wang, Cunxiang
Luo, Yu
Fan, Lin
Zhou, Yilin
Wang, Zikang
Gu, Xiaotao
Tang, Jie
Wang, Hongning
Huang, Minlie
author_facet Feng, Andrew Zhuoer
Wang, Cunxiang
Luo, Yu
Fan, Lin
Zhou, Yilin
Wang, Zikang
Gu, Xiaotao
Tang, Jie
Wang, Hongning
Huang, Minlie
contents Evaluating the writing capabilities of large language models (LLMs) remains a significant challenge due to the multidimensional nature of writing skills and the limitations of existing metrics. LLM's performance in thousand-words level and open-ended writing is inadequately assessed by traditional reference-based metrics or modern LLM-as-a-judge methods. We propose Tree-of-Writing (ToW), to resolve the implicit inconsistency often found when LLM-as-a-judge aggregates all sub-features in text evaluation. ToW incorporates a tree-structured workflow by explicitly modeling the aggregation weights of sub-features. We also present HowToBench, a large-scale Chinese writing benchmark encompassing 12 genres and 1302 instructions across three task categories: contextual completion, outline-guided writing, and open-ended generation. ToW successfully mitigates the biases, achieving a 0.93 Pearson correlation with human judgments. Furthermore, we detect that both overlap-based text generation metrics and popular LLM-as-a-judge practices are vulnerable to textual disturbances, while ToW is robust to them. We also uncover a negative correlation between input length and content-related scores in the Guide task, showcasing that it cannot be simply improved by input-side information piling.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19071
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
Feng, Andrew Zhuoer
Wang, Cunxiang
Luo, Yu
Fan, Lin
Zhou, Yilin
Wang, Zikang
Gu, Xiaotao
Tang, Jie
Wang, Hongning
Huang, Minlie
Computation and Language
Evaluating the writing capabilities of large language models (LLMs) remains a significant challenge due to the multidimensional nature of writing skills and the limitations of existing metrics. LLM's performance in thousand-words level and open-ended writing is inadequately assessed by traditional reference-based metrics or modern LLM-as-a-judge methods. We propose Tree-of-Writing (ToW), to resolve the implicit inconsistency often found when LLM-as-a-judge aggregates all sub-features in text evaluation. ToW incorporates a tree-structured workflow by explicitly modeling the aggregation weights of sub-features. We also present HowToBench, a large-scale Chinese writing benchmark encompassing 12 genres and 1302 instructions across three task categories: contextual completion, outline-guided writing, and open-ended generation. ToW successfully mitigates the biases, achieving a 0.93 Pearson correlation with human judgments. Furthermore, we detect that both overlap-based text generation metrics and popular LLM-as-a-judge practices are vulnerable to textual disturbances, while ToW is robust to them. We also uncover a negative correlation between input length and content-related scores in the Guide task, showcasing that it cannot be simply improved by input-side information piling.
title HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
topic Computation and Language
url https://arxiv.org/abs/2604.19071