Evaluating Text Style Transfer: A Nine-Language Benchmark for Text Detoxification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Protasov, Vitaly, Babakov, Nikolay, Dementieva, Daryna, Panchenko, Alexander
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908865762689024
author Protasov, Vitaly
Babakov, Nikolay
Dementieva, Daryna
Panchenko, Alexander
author_facet Protasov, Vitaly
Babakov, Nikolay
Dementieva, Daryna
Panchenko, Alexander
contents Despite notable advances in large language models (LLMs), reliable evaluation of text generation tasks such as text style transfer (TST) remains an open challenge. Existing research has shown that automatic metrics often correlate poorly with human judgments (Dementieva et al., 2024; Pauli et al., 2025), limiting our ability to assess model performance accurately. Furthermore, most prior work has focused primarily on English, while the evaluation of multilingual TST systems, particularly for text detoxification, remains largely underexplored. In this paper, we present the first comprehensive multilingual benchmarking study of evaluation metrics for text detoxification evaluation across nine languages: Arabic, Amharic, Chinese, English, German, Hindi, Russian, Spanish, and Ukrainian. Drawing inspiration from machine translation evaluation, we compare neural-based automatic metrics with LLM-as-a-judge approaches together with experiments on task-specific fine-tuned models. Our analysis reveals that the proposed metrics achieve significantly higher correlation with human judgments compared to baseline approaches. We also provide actionable insights and practical guidelines for building robust and reliable multilingual evaluation pipelines for text detoxification and related TST tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15557
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Text Style Transfer: A Nine-Language Benchmark for Text Detoxification
Protasov, Vitaly
Babakov, Nikolay
Dementieva, Daryna
Panchenko, Alexander
Computation and Language
Despite notable advances in large language models (LLMs), reliable evaluation of text generation tasks such as text style transfer (TST) remains an open challenge. Existing research has shown that automatic metrics often correlate poorly with human judgments (Dementieva et al., 2024; Pauli et al., 2025), limiting our ability to assess model performance accurately. Furthermore, most prior work has focused primarily on English, while the evaluation of multilingual TST systems, particularly for text detoxification, remains largely underexplored. In this paper, we present the first comprehensive multilingual benchmarking study of evaluation metrics for text detoxification evaluation across nine languages: Arabic, Amharic, Chinese, English, German, Hindi, Russian, Spanish, and Ukrainian. Drawing inspiration from machine translation evaluation, we compare neural-based automatic metrics with LLM-as-a-judge approaches together with experiments on task-specific fine-tuned models. Our analysis reveals that the proposed metrics achieve significantly higher correlation with human judgments compared to baseline approaches. We also provide actionable insights and practical guidelines for building robust and reliable multilingual evaluation pipelines for text detoxification and related TST tasks.
title Evaluating Text Style Transfer: A Nine-Language Benchmark for Text Detoxification
topic Computation and Language
url https://arxiv.org/abs/2507.15557