The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gómez-Rodríguez, Carlos, Williams, Paul
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911929978585088
author Gómez-Rodríguez, Carlos
Williams, Paul
author_facet Gómez-Rodríguez, Carlos
Williams, Paul
contents This is a summary of the paper "A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing", which was published in Findings of EMNLP 2023. We evaluate a range of recent state-of-the-art, instruction-tuned large language models (LLMs) on an English creative writing task, and compare them to human writers. For this purpose, we use a specifically-tailored prompt (based on an epic combat between Ignatius J. Reilly, main character of John Kennedy Toole's "A Confederacy of Dunces", and a pterodactyl) to minimize the risk of training data leakage and force the models to be creative rather than reusing existing stories. The same prompt is presented to LLMs and human writers, and evaluation is performed by humans using a detailed rubric including various aspects like fluency, style, originality or humor. Results show that some state-of-the-art commercial LLMs match or slightly outperform our human writers in most of the evaluated dimensions. Open-source LLMs lag behind. Humans keep a close lead in originality, and only the top three LLMs can handle humor at human-like levels.
format Preprint
id arxiv_https___arxiv_org_abs_2406_15891
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
Gómez-Rodríguez, Carlos
Williams, Paul
Computation and Language
68T50
I.2.7
This is a summary of the paper "A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing", which was published in Findings of EMNLP 2023. We evaluate a range of recent state-of-the-art, instruction-tuned large language models (LLMs) on an English creative writing task, and compare them to human writers. For this purpose, we use a specifically-tailored prompt (based on an epic combat between Ignatius J. Reilly, main character of John Kennedy Toole's "A Confederacy of Dunces", and a pterodactyl) to minimize the risk of training data leakage and force the models to be creative rather than reusing existing stories. The same prompt is presented to LLMs and human writers, and evaluation is performed by humans using a detailed rubric including various aspects like fluency, style, originality or humor. Results show that some state-of-the-art commercial LLMs match or slightly outperform our human writers in most of the evaluated dimensions. Open-source LLMs lag behind. Humans keep a close lead in originality, and only the top three LLMs can handle humor at human-like levels.
title The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
topic Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2406.15891