Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ying, Shuangshuang, Li, Yunwen, Qu, Xingwei, Li, Xin, Jin, Sheng, Liu, Minghao, Wen, Zhoufutu, Du, Xeron, Zheng, Tianyu, Zhang, Yichi, Ni, Letian, Cheng, Yuyang, Yang, Zhenzhu, Chen, Qiguang, Ding, Jingzhe, Long, Shengda, Zhou, Wangchunshu, Feng, Jiazhan, Zhong, Wanjun, Qin, Libo, Zhang, Ge, Huang, Wenhao, Che, Wanxiang, Lin, Chenghua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910002819629056
author Ying, Shuangshuang
Li, Yunwen
Qu, Xingwei
Li, Xin
Jin, Sheng
Liu, Minghao
Wen, Zhoufutu
Du, Xeron
Zheng, Tianyu
Zhang, Yichi
Ni, Letian
Cheng, Yuyang
Yang, Zhenzhu
Chen, Qiguang
Ding, Jingzhe
Long, Shengda
Zhou, Wangchunshu
Feng, Jiazhan
Zhong, Wanjun
Qin, Libo
Zhang, Ge
Huang, Wenhao
Che, Wanxiang
Lin, Chenghua
author_facet Ying, Shuangshuang
Li, Yunwen
Qu, Xingwei
Li, Xin
Jin, Sheng
Liu, Minghao
Wen, Zhoufutu
Du, Xeron
Zheng, Tianyu
Zhang, Yichi
Ni, Letian
Cheng, Yuyang
Yang, Zhenzhu
Chen, Qiguang
Ding, Jingzhe
Long, Shengda
Zhou, Wangchunshu
Feng, Jiazhan
Zhong, Wanjun
Qin, Libo
Zhang, Ge
Huang, Wenhao
Che, Wanxiang
Lin, Chenghua
contents Current preference learning methods achieve high accuracy on standard benchmarks but exhibit significant performance degradation when objective quality signals are removed. We introduce WritingPreferenceBench, a dataset of 1,800 human-annotated preference pairs (1,200 English, 600 Chinese) across 8 creative writing genres, where responses are matched for objective correctness, factual accuracy, and length. On this benchmark, sequence-based reward models--the standard architecture for RLHF--achieve only 52.7% mean accuracy, while zero-shot language model judges perform at 53.9%. In contrast, generative reward models that produce explicit reasoning chains achieve 81.8% accuracy. We observe high within-model variance across genres: individual models range from 18.2% to 81.8% accuracy across different writing categories, with standard deviations averaging 10.1%. This variance persists regardless of model scale, with 27B parameter models showing no consistent improvement over 8B variants. Our results suggest that current RLHF methods primarily learn to detect objective errors rather than capture subjective quality preferences (e.g., creativity, stylistic flair, and emotional resonance), and that successful preference modeling may require intermediate reasoning representations rather than direct classification.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14616
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
Ying, Shuangshuang
Li, Yunwen
Qu, Xingwei
Li, Xin
Jin, Sheng
Liu, Minghao
Wen, Zhoufutu
Du, Xeron
Zheng, Tianyu
Zhang, Yichi
Ni, Letian
Cheng, Yuyang
Yang, Zhenzhu
Chen, Qiguang
Ding, Jingzhe
Long, Shengda
Zhou, Wangchunshu
Feng, Jiazhan
Zhong, Wanjun
Qin, Libo
Zhang, Ge
Huang, Wenhao
Che, Wanxiang
Lin, Chenghua
Computation and Language
Artificial Intelligence
Current preference learning methods achieve high accuracy on standard benchmarks but exhibit significant performance degradation when objective quality signals are removed. We introduce WritingPreferenceBench, a dataset of 1,800 human-annotated preference pairs (1,200 English, 600 Chinese) across 8 creative writing genres, where responses are matched for objective correctness, factual accuracy, and length. On this benchmark, sequence-based reward models--the standard architecture for RLHF--achieve only 52.7% mean accuracy, while zero-shot language model judges perform at 53.9%. In contrast, generative reward models that produce explicit reasoning chains achieve 81.8% accuracy. We observe high within-model variance across genres: individual models range from 18.2% to 81.8% accuracy across different writing categories, with standard deviations averaging 10.1%. This variance persists regardless of model scale, with 27B parameter models showing no consistent improvement over 8B variants. Our results suggest that current RLHF methods primarily learn to detect objective errors rather than capture subjective quality preferences (e.g., creativity, stylistic flair, and emotional resonance), and that successful preference modeling may require intermediate reasoning representations rather than direct classification.
title Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.14616