HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Jingcong, Wan, Shijun, Wu, Xuehai, Li, Yitong, Chen, Qianglong, Tang, Duyu, Wang, Siyuan, Wei, Zhongyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914358107308032
author Liang, Jingcong
Wan, Shijun
Wu, Xuehai
Li, Yitong
Chen, Qianglong
Tang, Duyu
Wang, Siyuan
Wei, Zhongyu
author_facet Liang, Jingcong
Wan, Shijun
Wu, Xuehai
Li, Yitong
Chen, Qianglong
Tang, Duyu
Wang, Siyuan
Wei, Zhongyu
contents Large Reasoning Models (LRMs) have demonstrated impressive performance on complex tasks, including logical puzzle games that require deriving solutions satisfying all constraints. However, whether they can flexibly apply appropriate rules to varying conditions, particularly when faced with non-canonical game variants, remains an open question. Existing corpora focus on popular puzzles like 9x9 Sudoku, risking overfitting to canonical formats and memorization of solution patterns, which can mask deficiencies in understanding novel rules or adapting strategies to new variants. To address this, we introduce HardcoreLogic, a challenging benchmark of over 5,000 puzzles across 10 games, designed to test the robustness of LRMs on the "long-tail" of logical games. HardcoreLogic systematically transforms canonical puzzles through three dimensions: Increased Complexity (IC), Uncommon Elements (UE), and Unsolvable Puzzles (UP), reducing reliance on shortcut memorization. Evaluations on a diverse set of LRMs reveal significant performance drops, even for models achieving top scores on existing benchmarks, indicating heavy reliance on memorized stereotypes. While increased complexity is the dominant source of difficulty, models also struggle with subtle rule variations that do not necessarily increase puzzle difficulty. Our systematic error analysis on solvable and unsolvable puzzles further highlights gaps in genuine reasoning. Overall, HardcoreLogic exposes the limitations of current LRMs and establishes a benchmark for advancing high-level logical reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12563
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games
Liang, Jingcong
Wan, Shijun
Wu, Xuehai
Li, Yitong
Chen, Qianglong
Tang, Duyu
Wang, Siyuan
Wei, Zhongyu
Artificial Intelligence
Large Reasoning Models (LRMs) have demonstrated impressive performance on complex tasks, including logical puzzle games that require deriving solutions satisfying all constraints. However, whether they can flexibly apply appropriate rules to varying conditions, particularly when faced with non-canonical game variants, remains an open question. Existing corpora focus on popular puzzles like 9x9 Sudoku, risking overfitting to canonical formats and memorization of solution patterns, which can mask deficiencies in understanding novel rules or adapting strategies to new variants. To address this, we introduce HardcoreLogic, a challenging benchmark of over 5,000 puzzles across 10 games, designed to test the robustness of LRMs on the "long-tail" of logical games. HardcoreLogic systematically transforms canonical puzzles through three dimensions: Increased Complexity (IC), Uncommon Elements (UE), and Unsolvable Puzzles (UP), reducing reliance on shortcut memorization. Evaluations on a diverse set of LRMs reveal significant performance drops, even for models achieving top scores on existing benchmarks, indicating heavy reliance on memorized stereotypes. While increased complexity is the dominant source of difficulty, models also struggle with subtle rule variations that do not necessarily increase puzzle difficulty. Our systematic error analysis on solvable and unsolvable puzzles further highlights gaps in genuine reasoning. Overall, HardcoreLogic exposes the limitations of current LRMs and establishes a benchmark for advancing high-level logical reasoning.
title HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games
topic Artificial Intelligence
url https://arxiv.org/abs/2510.12563