What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Mengtao, Wu, Sifan, Zhang, Huan, Sima, Qi, Liu, Bang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can LLMs Ask Good Questions?
von: Zhang, Yueheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yueheng, et al.
Veröffentlicht: (2025)
Teaching LLMs to Ask: Self-Querying Category-Theoretic Planning for Under-Specified Reasoning
von: Qu, Shuhui
Veröffentlicht: (2026)
von: Qu, Shuhui
Veröffentlicht: (2026)
What Is Next for LLMs? Next-Generation AI Computing Hardware Using Photonic Chips
von: Li, Renjie, et al.
Veröffentlicht: (2025)
von: Li, Renjie, et al.
Veröffentlicht: (2025)
Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles
von: Chen, Qi, et al.
Veröffentlicht: (2024)
von: Chen, Qi, et al.
Veröffentlicht: (2024)
What Would Happen Next? Predicting Consequences from An Event Causality Graph
von: Zhan, Chuanhong, et al.
Veröffentlicht: (2024)
von: Zhan, Chuanhong, et al.
Veröffentlicht: (2024)
Ask Good Questions for Large Language Models
von: Wu, Qi, et al.
Veröffentlicht: (2025)
von: Wu, Qi, et al.
Veröffentlicht: (2025)
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
von: Long, Yitao, et al.
Veröffentlicht: (2025)
von: Long, Yitao, et al.
Veröffentlicht: (2025)
Improving Clinical Note Generation from Complex Doctor-Patient Conversation
von: Li, Yizhan, et al.
Veröffentlicht: (2024)
von: Li, Yizhan, et al.
Veröffentlicht: (2024)
AgentAsk: Multi-Agent Systems Need to Ask
von: Lin, Bohan, et al.
Veröffentlicht: (2025)
von: Lin, Bohan, et al.
Veröffentlicht: (2025)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
von: Tyagi, Nemika, et al.
Veröffentlicht: (2024)
von: Tyagi, Nemika, et al.
Veröffentlicht: (2024)
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
Bone Soups: A Seek-and-Soup Model Merging Approach for Controllable Multi-Objective Generation
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
Fail Fast, or Ask: Mitigating the Deficiencies of Reasoning LLMs with Human-in-the-Loop Systems Engineering
von: Zellinger, Michael J., et al.
Veröffentlicht: (2025)
von: Zellinger, Michael J., et al.
Veröffentlicht: (2025)
CircuitSeer: Mining High-Quality Data by Probing Mathematical Reasoning Circuits in LLMs
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning
von: Ramrakhya, Ram, et al.
Veröffentlicht: (2025)
von: Ramrakhya, Ram, et al.
Veröffentlicht: (2025)
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
Students' Feedback Requests and Interactions with the SCRIPT Chatbot: Do They Get What They Ask For?
von: Scholl, Andreas, et al.
Veröffentlicht: (2025)
von: Scholl, Andreas, et al.
Veröffentlicht: (2025)
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models
von: Lyu, Zesen, et al.
Veröffentlicht: (2025)
von: Lyu, Zesen, et al.
Veröffentlicht: (2025)
SoupLM: Model Integration in Large Language and Multi-Modal Models
von: Bai, Yue, et al.
Veröffentlicht: (2024)
von: Bai, Yue, et al.
Veröffentlicht: (2024)
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
von: Tang, Kexian, et al.
Veröffentlicht: (2025)
von: Tang, Kexian, et al.
Veröffentlicht: (2025)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
von: Shin, Jisu, et al.
Veröffentlicht: (2024)
von: Shin, Jisu, et al.
Veröffentlicht: (2024)
VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models
von: Xu, Hefei, et al.
Veröffentlicht: (2026)
von: Xu, Hefei, et al.
Veröffentlicht: (2026)
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
von: Şenol, Ali, et al.
Veröffentlicht: (2026)
von: Şenol, Ali, et al.
Veröffentlicht: (2026)
The Token Games: Evaluating Language Model Reasoning with Puzzle Duels
von: Henniger, Simon, et al.
Veröffentlicht: (2026)
von: Henniger, Simon, et al.
Veröffentlicht: (2026)
PuzzleJAX: A Benchmark for Reasoning and Learning
von: Earle, Sam, et al.
Veröffentlicht: (2025)
von: Earle, Sam, et al.
Veröffentlicht: (2025)
Measuring Iterative Temporal Reasoning with Time Puzzles
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2026)
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2026)
HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games
von: Liang, Jingcong, et al.
Veröffentlicht: (2025)
von: Liang, Jingcong, et al.
Veröffentlicht: (2025)
DeepImagine: Learning Biomedical Reasoning via Successive Counterfactual Imagining
von: Zheng, Youze, et al.
Veröffentlicht: (2026)
von: Zheng, Youze, et al.
Veröffentlicht: (2026)
Stepwise Reasoning Error Disruption Attack of LLMs
von: Peng, Jingyu, et al.
Veröffentlicht: (2024)
von: Peng, Jingyu, et al.
Veröffentlicht: (2024)
RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
von: Jin, Hangzhan, et al.
Veröffentlicht: (2025)
von: Jin, Hangzhan, et al.
Veröffentlicht: (2025)
PUZZLED: Jailbreaking LLMs through Word-Based Puzzles
von: Ahn, Yelim, et al.
Veröffentlicht: (2025)
von: Ahn, Yelim, et al.
Veröffentlicht: (2025)
DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities
von: Zhuang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhuang, Tianyi, et al.
Veröffentlicht: (2025)
What You See is What You Ask: Evaluating Audio Descriptions
von: Kala, Divy, et al.
Veröffentlicht: (2025)
von: Kala, Divy, et al.
Veröffentlicht: (2025)
Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
von: Zhao, Yuze, et al.
Veröffentlicht: (2026)
von: Zhao, Yuze, et al.
Veröffentlicht: (2026)
Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
von: Ghosal, Deepanway, et al.
Veröffentlicht: (2024)
von: Ghosal, Deepanway, et al.
Veröffentlicht: (2024)
Frog Soup: Zero-Shot, In-Context, and Sample-Efficient Frogger Agents
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
M^4olGen: Multi-Agent, Multi-Stage Molecular Generation under Precise Multi-Property Constraints
von: Li, Yizhan, et al.
Veröffentlicht: (2026)
von: Li, Yizhan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Can LLMs Ask Good Questions?
von: Zhang, Yueheng, et al.
Veröffentlicht: (2025) -
Teaching LLMs to Ask: Self-Querying Category-Theoretic Planning for Under-Specified Reasoning
von: Qu, Shuhui
Veröffentlicht: (2026) -
What Is Next for LLMs? Next-Generation AI Computing Hardware Using Photonic Chips
von: Li, Renjie, et al.
Veröffentlicht: (2025) -
Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles
von: Chen, Qi, et al.
Veröffentlicht: (2024) -
What Would Happen Next? Predicting Consequences from An Event Causality Graph
von: Zhan, Chuanhong, et al.
Veröffentlicht: (2024)