Art or Artifice? Large Language Models and the False Promise of Creativity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chakrabarty, Tuhin, Laban, Philippe, Agarwal, Divyansh, Muresan, Smaranda, Wu, Chien-Sheng
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916151190093824
author Chakrabarty, Tuhin
Laban, Philippe
Agarwal, Divyansh
Muresan, Smaranda
Wu, Chien-Sheng
author_facet Chakrabarty, Tuhin
Laban, Philippe
Agarwal, Divyansh
Muresan, Smaranda
Wu, Chien-Sheng
contents Researchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by the Torrance Test of Creative Thinking (TTCT), which measures creativity as a process, we use the Consensual Assessment Technique [3] and propose the Torrance Test of Creative Writing (TTCW) to evaluate creativity as a product. TTCW consists of 14 binary tests organized into the original dimensions of Fluency, Flexibility, Originality, and Elaboration. We recruit 10 creative writers and implement a human assessment of 48 stories written either by professional authors or LLMs using TTCW. Our analysis shows that LLM-generated stories pass 3-10X less TTCW tests than stories written by professionals. In addition, we explore the use of LLMs as assessors to automate the TTCW evaluation, revealing that none of the LLMs positively correlate with the expert assessments.
format Preprint
id arxiv_https___arxiv_org_abs_2309_14556
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Art or Artifice? Large Language Models and the False Promise of Creativity
Chakrabarty, Tuhin
Laban, Philippe
Agarwal, Divyansh
Muresan, Smaranda
Wu, Chien-Sheng
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Researchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by the Torrance Test of Creative Thinking (TTCT), which measures creativity as a process, we use the Consensual Assessment Technique [3] and propose the Torrance Test of Creative Writing (TTCW) to evaluate creativity as a product. TTCW consists of 14 binary tests organized into the original dimensions of Fluency, Flexibility, Originality, and Elaboration. We recruit 10 creative writers and implement a human assessment of 48 stories written either by professional authors or LLMs using TTCW. Our analysis shows that LLM-generated stories pass 3-10X less TTCW tests than stories written by professionals. In addition, we explore the use of LLMs as assessors to automate the TTCW evaluation, revealing that none of the LLMs positively correlate with the expert assessments.
title Art or Artifice? Large Language Models and the False Promise of Creativity
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2309.14556