CEval: A Benchmark for Evaluating Counterfactual Text Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Van Bach, Schlötterer, Jörg, Seifert, Christin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917747307315200
author Nguyen, Van Bach
Schlötterer, Jörg
Seifert, Christin
author_facet Nguyen, Van Bach
Schlötterer, Jörg
Seifert, Christin
contents Counterfactual text generation aims to minimally change a text, such that it is classified differently. Judging advancements in method development for counterfactual text generation is hindered by a non-uniform usage of data sets and metrics in related work. We propose CEval, a benchmark for comparing counterfactual text generation methods. CEval unifies counterfactual and text quality metrics, includes common counterfactual datasets with human annotations, standard baselines (MICE, GDBA, CREST) and the open-source language model LLAMA-2. Our experiments found no perfect method for generating counterfactual text. Methods that excel at counterfactual metrics often produce lower-quality text while LLMs with simple prompts generate high-quality text but struggle with counterfactual criteria. By making CEval available as an open-source Python library, we encourage the community to contribute more methods and maintain consistent evaluation in future work.
format Preprint
id arxiv_https___arxiv_org_abs_2404_17475
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CEval: A Benchmark for Evaluating Counterfactual Text Generation
Nguyen, Van Bach
Schlötterer, Jörg
Seifert, Christin
Computation and Language
Artificial Intelligence
Counterfactual text generation aims to minimally change a text, such that it is classified differently. Judging advancements in method development for counterfactual text generation is hindered by a non-uniform usage of data sets and metrics in related work. We propose CEval, a benchmark for comparing counterfactual text generation methods. CEval unifies counterfactual and text quality metrics, includes common counterfactual datasets with human annotations, standard baselines (MICE, GDBA, CREST) and the open-source language model LLAMA-2. Our experiments found no perfect method for generating counterfactual text. Methods that excel at counterfactual metrics often produce lower-quality text while LLMs with simple prompts generate high-quality text but struggle with counterfactual criteria. By making CEval available as an open-source Python library, we encourage the community to contribute more methods and maintain consistent evaluation in future work.
title CEval: A Benchmark for Evaluating Counterfactual Text Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.17475