Do LLMs Agree on the Creativity Evaluation of Alternative Uses?
Fuente:
arXiv
Saved in:
| Main Authors: | Rabeyah, Abdullah Al, Góes, Fabrício, Volpe, Marco, Medeiros, Talles |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are Frontier Large Language Models Suitable for Q&A in Science Centres?
by: Watson, Jacob, et al.
Published: (2024)
by: Watson, Jacob, et al.
Published: (2024)
Do LLMs Dream of Ontologies?
by: Bombieri, Marco, et al.
Published: (2024)
by: Bombieri, Marco, et al.
Published: (2024)
We Argue to Agree: Towards Personality-Driven Argumentation-Based Negotiation Dialogue Systems for Tourism
by: Priya, Priyanshu, et al.
Published: (2025)
by: Priya, Priyanshu, et al.
Published: (2025)
Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs
by: Marco, Guillermo, et al.
Published: (2024)
by: Marco, Guillermo, et al.
Published: (2024)
BengaliFig: A Low-Resource Challenge for Figurative and Culturally Grounded Reasoning in Bengali
by: Sefat, Abdullah Al
Published: (2025)
by: Sefat, Abdullah Al
Published: (2025)
CreativityPrism: A Holistic Evaluation Framework for Large Language Model Creativity
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
Assessing the Creativity of LLMs in Proposing Novel Solutions to Mathematical Problems
by: Ye, Junyi, et al.
Published: (2024)
by: Ye, Junyi, et al.
Published: (2024)
Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
by: Amirizaniani, Maryam, et al.
Published: (2024)
by: Amirizaniani, Maryam, et al.
Published: (2024)
Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs
by: Banerjee, Mohor, et al.
Published: (2025)
by: Banerjee, Mohor, et al.
Published: (2025)
Evaluating LLMs on Entity Disambiguation in Tables
by: Belotti, Federico, et al.
Published: (2024)
by: Belotti, Federico, et al.
Published: (2024)
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
Steering Large Language Models to Evaluate and Amplify Creativity
by: Olson, Matthew Lyle, et al.
Published: (2024)
by: Olson, Matthew Lyle, et al.
Published: (2024)
How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs?
by: Doostmohammadi, Ehsan, et al.
Published: (2024)
by: Doostmohammadi, Ehsan, et al.
Published: (2024)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs
by: Guo, Yanzhu, et al.
Published: (2024)
by: Guo, Yanzhu, et al.
Published: (2024)
Do LLMs have Consistent Values?
by: Rozen, Naama, et al.
Published: (2024)
by: Rozen, Naama, et al.
Published: (2024)
How Do LLMs Use Their Depth?
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
A Robot Walks into a Bar: Can Language Models Serve as Creativity Support Tools for Comedy? An Evaluation of LLMs' Humour Alignment with Comedians
by: Mirowski, Piotr Wojciech, et al.
Published: (2024)
by: Mirowski, Piotr Wojciech, et al.
Published: (2024)
Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs
by: Camassa, Carolina, et al.
Published: (2026)
by: Camassa, Carolina, et al.
Published: (2026)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
by: Mustapha, Ahmad, et al.
Published: (2024)
by: Mustapha, Ahmad, et al.
Published: (2024)
On the Suitability of pre-trained foundational LLMs for Analysis in German Legal Education
by: Wendlinger, Lorenz, et al.
Published: (2024)
by: Wendlinger, Lorenz, et al.
Published: (2024)
Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark
by: Choi, Minje, et al.
Published: (2023)
by: Choi, Minje, et al.
Published: (2023)
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?
by: Li, Senyu, et al.
Published: (2025)
by: Li, Senyu, et al.
Published: (2025)
From Words to Proverbs: Evaluating LLMs Linguistic and Cultural Competence in Saudi Dialects with Absher
by: Al-Monef, Renad, et al.
Published: (2025)
by: Al-Monef, Renad, et al.
Published: (2025)
LLMs Do Not Grade Essays Like Humans
by: Mathew, Jerin George, et al.
Published: (2026)
by: Mathew, Jerin George, et al.
Published: (2026)
Do LLMs Benefit From Their Own Words?
by: Huang, Jenny Y., et al.
Published: (2026)
by: Huang, Jenny Y., et al.
Published: (2026)
Evaluating Creative Short Story Generation in Humans and Large Language Models
by: Ismayilzada, Mete, et al.
Published: (2024)
by: Ismayilzada, Mete, et al.
Published: (2024)
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
by: Fein, Daniel, et al.
Published: (2025)
by: Fein, Daniel, et al.
Published: (2025)
Beyond Natural Language: LLMs Leveraging Alternative Formats for Enhanced Reasoning and Communication
by: Chen, Weize, et al.
Published: (2024)
by: Chen, Weize, et al.
Published: (2024)
Do LLMs Align with My Task? Evaluating Text-to-SQL via Dataset Alignment
by: Rafiei, Davood, et al.
Published: (2025)
by: Rafiei, Davood, et al.
Published: (2025)
Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification
by: Demir, M. Mikail, et al.
Published: (2026)
by: Demir, M. Mikail, et al.
Published: (2026)
IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
by: Elmahjub, Ezieddin, et al.
Published: (2026)
by: Elmahjub, Ezieddin, et al.
Published: (2026)
Serendipity by Design: Evaluating the Impact of Cross-domain Mappings on Human and LLM Creativity
by: Liu, Qiawen Ella, et al.
Published: (2026)
by: Liu, Qiawen Ella, et al.
Published: (2026)
Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach
by: Li, Ruizhe, et al.
Published: (2025)
by: Li, Ruizhe, et al.
Published: (2025)
Do LLMs estimate uncertainty well in instruction-following?
by: Heo, Juyeon, et al.
Published: (2024)
by: Heo, Juyeon, et al.
Published: (2024)
Do LLMs "know" internally when they follow instructions?
by: Heo, Juyeon, et al.
Published: (2024)
by: Heo, Juyeon, et al.
Published: (2024)
Do LLMs "Feel"? Emotion Circuits Discovery and Control
by: Wang, Chenxi, et al.
Published: (2025)
by: Wang, Chenxi, et al.
Published: (2025)
How Well Do LLMs Understand Tunisian Arabic?
by: Mahdi, Mohamed
Published: (2025)
by: Mahdi, Mohamed
Published: (2025)
CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing
by: Qian, Cheng, et al.
Published: (2026)
by: Qian, Cheng, et al.
Published: (2026)
Bridging the Gap in Bangla Healthcare: Machine Learning Based Disease Prediction Using a Symptoms-Disease Dataset
by: Zannat, Rowzatul, et al.
Published: (2026)
by: Zannat, Rowzatul, et al.
Published: (2026)
Similar Items
-
Are Frontier Large Language Models Suitable for Q&A in Science Centres?
by: Watson, Jacob, et al.
Published: (2024) -
Do LLMs Dream of Ontologies?
by: Bombieri, Marco, et al.
Published: (2024) -
We Argue to Agree: Towards Personality-Driven Argumentation-Based Negotiation Dialogue Systems for Tourism
by: Priya, Priyanshu, et al.
Published: (2025) -
Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs
by: Marco, Guillermo, et al.
Published: (2024) -
BengaliFig: A Low-Resource Challenge for Figurative and Culturally Grounded Reasoning in Bengali
by: Sefat, Abdullah Al
Published: (2025)