Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Simeng, Dai, Howard, Xia, Stephen, Zhang, Grant, Liu, Chen, Chen, Lichang, Nguyen, Hoang Huy, Mei, Hongyuan, Mao, Jiayuan, McCoy, R. Thomas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908617702113280
author Han, Simeng
Dai, Howard
Xia, Stephen
Zhang, Grant
Liu, Chen
Chen, Lichang
Nguyen, Hoang Huy
Mei, Hongyuan
Mao, Jiayuan
McCoy, R. Thomas
author_facet Han, Simeng
Dai, Howard
Xia, Stephen
Zhang, Grant
Liu, Chen
Chen, Lichang
Nguyen, Hoang Huy
Mei, Hongyuan
Mao, Jiayuan
McCoy, R. Thomas
contents Accuracy remains a standard metric for evaluating AI systems, but it offers limited insight into how models arrive at their solutions. In this work, we introduce a benchmark based on brainteasers written in long narrative form to probe more deeply into the types of reasoning strategies that models use. Brainteasers are well-suited for this goal because they can be solved with multiple approaches, such as a few-step solution that uses a creative insight or a longer solution that uses more brute force. We investigate large language models (LLMs) across multiple layers of reasoning, focusing not only on correctness but also on the quality and creativity of their solutions. We investigate many aspects of the reasoning process: (1) semantic parsing of the brainteasers into precise mathematical competition style formats; (2) generating solutions from these mathematical forms; (3) self-correcting solutions based on gold solutions; (4) producing step-by-step sketches of solutions; and (5) making use of hints. We find that LLMs are in many cases able to find creative, insightful solutions to brainteasers, suggesting that they capture some of the capacities needed to solve novel problems in creative ways. Nonetheless, there also remain situations where they rely on brute force despite the availability of more efficient, creative solutions, highlighting a potential direction for improvement in the reasoning abilities of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10844
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models
Han, Simeng
Dai, Howard
Xia, Stephen
Zhang, Grant
Liu, Chen
Chen, Lichang
Nguyen, Hoang Huy
Mei, Hongyuan
Mao, Jiayuan
McCoy, R. Thomas
Artificial Intelligence
Computation and Language
Accuracy remains a standard metric for evaluating AI systems, but it offers limited insight into how models arrive at their solutions. In this work, we introduce a benchmark based on brainteasers written in long narrative form to probe more deeply into the types of reasoning strategies that models use. Brainteasers are well-suited for this goal because they can be solved with multiple approaches, such as a few-step solution that uses a creative insight or a longer solution that uses more brute force. We investigate large language models (LLMs) across multiple layers of reasoning, focusing not only on correctness but also on the quality and creativity of their solutions. We investigate many aspects of the reasoning process: (1) semantic parsing of the brainteasers into precise mathematical competition style formats; (2) generating solutions from these mathematical forms; (3) self-correcting solutions based on gold solutions; (4) producing step-by-step sketches of solutions; and (5) making use of hints. We find that LLMs are in many cases able to find creative, insightful solutions to brainteasers, suggesting that they capture some of the capacities needed to solve novel problems in creative ways. Nonetheless, there also remain situations where they rely on brute force despite the availability of more efficient, creative solutions, highlighting a potential direction for improvement in the reasoning abilities of LLMs.
title Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.10844