PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Long, Yitao, Jiang, Yuru, Liu, Hongjun, Zhao, Yilun, Sun, Jingchen, Shen, Yiqiu, Zhao, Chen, Cohan, Arman, Shasha, Dennis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914080268222464
author Long, Yitao
Jiang, Yuru
Liu, Hongjun
Zhao, Yilun
Sun, Jingchen
Shen, Yiqiu
Zhao, Chen
Cohan, Arman
Shasha, Dennis
author_facet Long, Yitao
Jiang, Yuru
Liu, Hongjun
Zhao, Yilun
Sun, Jingchen
Shen, Yiqiu
Zhao, Chen
Cohan, Arman
Shasha, Dennis
contents This work investigates the reasoning and planning capabilities of foundation models and their scalability in complex, dynamic environments. We introduce PuzzlePlex, a benchmark designed to assess these capabilities through a diverse set of puzzles. PuzzlePlex consists of 15 types of puzzles, including deterministic and stochastic games of varying difficulty, as well as single-player and two-player scenarios. The PuzzlePlex framework provides a comprehensive environment for each game, and supports extensibility to generate more challenging instances as foundation models evolve. Additionally, we implement customized game-playing strategies for comparison. Building on this benchmark, we develop fine-grained metrics to measure performance and conduct an in-depth analysis of frontier foundation models across two settings: instruction-based and code-based. Furthermore, we systematically investigate their scaling limits. Our findings show that reasoning models outperform others in instruction-based settings, while code-based execution presents greater challenges but offers a scalable and efficient alternative. PuzzlePlex enables targeted evaluation and guides future improvements in reasoning, planning, and generalization for foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06475
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
Long, Yitao
Jiang, Yuru
Liu, Hongjun
Zhao, Yilun
Sun, Jingchen
Shen, Yiqiu
Zhao, Chen
Cohan, Arman
Shasha, Dennis
Artificial Intelligence
Computation and Language
This work investigates the reasoning and planning capabilities of foundation models and their scalability in complex, dynamic environments. We introduce PuzzlePlex, a benchmark designed to assess these capabilities through a diverse set of puzzles. PuzzlePlex consists of 15 types of puzzles, including deterministic and stochastic games of varying difficulty, as well as single-player and two-player scenarios. The PuzzlePlex framework provides a comprehensive environment for each game, and supports extensibility to generate more challenging instances as foundation models evolve. Additionally, we implement customized game-playing strategies for comparison. Building on this benchmark, we develop fine-grained metrics to measure performance and conduct an in-depth analysis of frontier foundation models across two settings: instruction-based and code-based. Furthermore, we systematically investigate their scaling limits. Our findings show that reasoning models outperform others in instruction-based settings, while code-based execution presents greater challenges but offers a scalable and efficient alternative. PuzzlePlex enables targeted evaluation and guides future improvements in reasoning, planning, and generalization for foundation models.
title PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.06475