Saved in:
Bibliographic Details
Main Authors: Zhou, Mengtao, Wu, Sifan, Zhang, Huan, Sima, Qi, Liu, Bang
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.10358
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912537424953344
author Zhou, Mengtao
Wu, Sifan
Zhang, Huan
Sima, Qi
Liu, Bang
author_facet Zhou, Mengtao
Wu, Sifan
Zhang, Huan
Sima, Qi
Liu, Bang
contents We investigate the capacity of Large Language Models (LLMs) for imaginative reasoning--the proactive construction, testing, and revision of hypotheses in information-sparse environments. Existing benchmarks, often static or focused on social deduction, fail to capture the dynamic, exploratory nature of this reasoning process. To address this gap, we introduce a comprehensive research framework based on the classic "Turtle Soup" game, integrating a benchmark, an agent, and an evaluation protocol. We present TurtleSoup-Bench, the first large-scale, bilingual, interactive benchmark for imaginative reasoning, comprising 800 turtle soup puzzles sourced from both the Internet and expert authors. We also propose Mosaic-Agent, a novel agent designed to assess LLMs' performance in this setting. To evaluate reasoning quality, we develop a multi-dimensional protocol measuring logical consistency, detail completion, and conclusion alignment. Experiments with leading LLMs reveal clear capability limits, common failure patterns, and a significant performance gap compared to humans. Our work offers new insights into LLMs' imaginative reasoning and establishes a foundation for future research on exploratory agent behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10358
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
Zhou, Mengtao
Wu, Sifan
Zhang, Huan
Sima, Qi
Liu, Bang
Artificial Intelligence
We investigate the capacity of Large Language Models (LLMs) for imaginative reasoning--the proactive construction, testing, and revision of hypotheses in information-sparse environments. Existing benchmarks, often static or focused on social deduction, fail to capture the dynamic, exploratory nature of this reasoning process. To address this gap, we introduce a comprehensive research framework based on the classic "Turtle Soup" game, integrating a benchmark, an agent, and an evaluation protocol. We present TurtleSoup-Bench, the first large-scale, bilingual, interactive benchmark for imaginative reasoning, comprising 800 turtle soup puzzles sourced from both the Internet and expert authors. We also propose Mosaic-Agent, a novel agent designed to assess LLMs' performance in this setting. To evaluate reasoning quality, we develop a multi-dimensional protocol measuring logical consistency, detail completion, and conclusion alignment. Experiments with leading LLMs reveal clear capability limits, common failure patterns, and a significant performance gap compared to humans. Our work offers new insights into LLMs' imaginative reasoning and establishes a foundation for future research on exploratory agent behavior.
title What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
topic Artificial Intelligence
url https://arxiv.org/abs/2508.10358