Saved in:
Bibliographic Details
Main Authors: Jangra, Anubhav, Sarrafzadeh, Bahareh, Cucerzan, Silviu, de Wynter, Adrian, Jauhar, Sujay Kumar
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.06374
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918160731471872
author Jangra, Anubhav
Sarrafzadeh, Bahareh
Cucerzan, Silviu
de Wynter, Adrian
Jauhar, Sujay Kumar
author_facet Jangra, Anubhav
Sarrafzadeh, Bahareh
Cucerzan, Silviu
de Wynter, Adrian
Jauhar, Sujay Kumar
contents With the surge of large language models (LLMs) and their ability to produce customized output, style-personalized text generation--"write like me"--has become a rapidly growing area of interest. However, style personalization is highly specific, relative to every user, and depends strongly on the pragmatic context, which makes it uniquely challenging. Although prior research has introduced benchmarks and metrics for this area, they tend to be non-standardized and have known limitations (e.g., poor correlation with human subjects). LLMs have been found to not capture author-specific style well, it follows that the metrics themselves must be scrutinized carefully. In this work we critically examine the effectiveness of the most common metrics used in the field, such as BLEU, embeddings, and LLMs-as-judges. We evaluate these metrics using our proposed style discrimination benchmark, which spans eight diverse writing tasks across three evaluation settings: domain discrimination, authorship attribution, and LLM-generated personalized vs non-personalized discrimination. We find strong evidence that employing ensembles of diverse evaluation metrics consistently outperforms single-evaluator methods, and conclude by providing guidance on how to reliably assess style-personalized text generation.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06374
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Style-Personalized Text Generation: Challenges and Directions
Jangra, Anubhav
Sarrafzadeh, Bahareh
Cucerzan, Silviu
de Wynter, Adrian
Jauhar, Sujay Kumar
Computation and Language
With the surge of large language models (LLMs) and their ability to produce customized output, style-personalized text generation--"write like me"--has become a rapidly growing area of interest. However, style personalization is highly specific, relative to every user, and depends strongly on the pragmatic context, which makes it uniquely challenging. Although prior research has introduced benchmarks and metrics for this area, they tend to be non-standardized and have known limitations (e.g., poor correlation with human subjects). LLMs have been found to not capture author-specific style well, it follows that the metrics themselves must be scrutinized carefully. In this work we critically examine the effectiveness of the most common metrics used in the field, such as BLEU, embeddings, and LLMs-as-judges. We evaluate these metrics using our proposed style discrimination benchmark, which spans eight diverse writing tasks across three evaluation settings: domain discrimination, authorship attribution, and LLM-generated personalized vs non-personalized discrimination. We find strong evidence that employing ensembles of diverse evaluation metrics consistently outperforms single-evaluator methods, and conclude by providing guidance on how to reliably assess style-personalized text generation.
title Evaluating Style-Personalized Text Generation: Challenges and Directions
topic Computation and Language
url https://arxiv.org/abs/2508.06374