Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Ziliang, Hu, Renfen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912645952569344
author Qiu, Ziliang
Hu, Renfen
author_facet Qiu, Ziliang
Hu, Renfen
contents The evaluation of LLMs' creativity represents a crucial research domain, though challenges such as data contamination and costly human assessments often impede progress. Drawing inspiration from human creativity assessment, we propose PACE, asking LLMs to generate Parallel Association Chains to Evaluate their creativity. PACE minimizes the risk of data contamination and offers a straightforward, highly efficient evaluation, as evidenced by its strong correlation with Chatbot Arena Creative Writing rankings (Spearman's $ρ= 0.739$, $p < 0.001$) across various proprietary and open-source models. A comparative analysis of associative creativity between LLMs and humans reveals that while high-performing LLMs achieve scores comparable to average human performance, professional humans consistently outperform LLMs. Furthermore, linguistic analysis reveals that both humans and LLMs exhibit a trend of decreasing concreteness in their associations, and humans demonstrating a greater diversity of associative patterns.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12110
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models
Qiu, Ziliang
Hu, Renfen
Computation and Language
Artificial Intelligence
The evaluation of LLMs' creativity represents a crucial research domain, though challenges such as data contamination and costly human assessments often impede progress. Drawing inspiration from human creativity assessment, we propose PACE, asking LLMs to generate Parallel Association Chains to Evaluate their creativity. PACE minimizes the risk of data contamination and offers a straightforward, highly efficient evaluation, as evidenced by its strong correlation with Chatbot Arena Creative Writing rankings (Spearman's $ρ= 0.739$, $p < 0.001$) across various proprietary and open-source models. A comparative analysis of associative creativity between LLMs and humans reveals that while high-performing LLMs achieve scores comparable to average human performance, professional humans consistently outperform LLMs. Furthermore, linguistic analysis reveals that both humans and LLMs exhibit a trend of decreasing concreteness in their associations, and humans demonstrating a greater diversity of associative patterns.
title Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.12110