Evaluating Creative Short Story Generation in Humans and Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ismayilzada, Mete, Stevenson, Claire, van der Plas, Lonneke
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910936512593920
author Ismayilzada, Mete
Stevenson, Claire
van der Plas, Lonneke
author_facet Ismayilzada, Mete
Stevenson, Claire
van der Plas, Lonneke
contents Story-writing is a fundamental aspect of human imagination, relying heavily on creativity to produce narratives that are novel, effective, and surprising. While large language models (LLMs) have demonstrated the ability to generate high-quality stories, their creative story-writing capabilities remain under-explored. In this work, we conduct a systematic analysis of creativity in short story generation across 60 LLMs and 60 people using a five-sentence cue-word-based creative story-writing task. We use measures to automatically evaluate model- and human-generated stories across several dimensions of creativity, including novelty, surprise, diversity, and linguistic complexity. We also collect creativity ratings and Turing Test classifications from non-expert and expert human raters and LLMs. Automated metrics show that LLMs generate stylistically complex stories, but tend to fall short in terms of novelty, surprise and diversity when compared to average human writers. Expert ratings generally coincide with automated metrics. However, LLMs and non-experts rate LLM stories to be more creative than human-generated stories. We discuss why and how these differences in ratings occur, and their implications for both human and artificial creativity.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02316
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Creative Short Story Generation in Humans and Large Language Models
Ismayilzada, Mete
Stevenson, Claire
van der Plas, Lonneke
Computation and Language
Artificial Intelligence
Story-writing is a fundamental aspect of human imagination, relying heavily on creativity to produce narratives that are novel, effective, and surprising. While large language models (LLMs) have demonstrated the ability to generate high-quality stories, their creative story-writing capabilities remain under-explored. In this work, we conduct a systematic analysis of creativity in short story generation across 60 LLMs and 60 people using a five-sentence cue-word-based creative story-writing task. We use measures to automatically evaluate model- and human-generated stories across several dimensions of creativity, including novelty, surprise, diversity, and linguistic complexity. We also collect creativity ratings and Turing Test classifications from non-expert and expert human raters and LLMs. Automated metrics show that LLMs generate stylistically complex stories, but tend to fall short in terms of novelty, surprise and diversity when compared to average human writers. Expert ratings generally coincide with automated metrics. However, LLMs and non-experts rate LLM stories to be more creative than human-generated stories. We discuss why and how these differences in ratings occur, and their implications for both human and artificial creativity.
title Evaluating Creative Short Story Generation in Humans and Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.02316