PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Qian, Yusu, Wan, Cheng, Jia, Chao, Yang, Yinfei, Zhao, Qingyu, Gan, Zhe
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909936488808448
author Qian, Yusu
Wan, Cheng
Jia, Chao
Yang, Yinfei
Zhao, Qingyu
Gan, Zhe
author_facet Qian, Yusu
Wan, Cheng
Jia, Chao
Yang, Yinfei
Zhao, Qingyu
Gan, Zhe
contents Multimodal large language models (MLLMs) have achieved remarkable progress on vision-language tasks, yet their reasoning processes remain sometimes unreliable. We introduce PRISM-Bench, a benchmark of puzzle-based visual challenges designed to evaluate not only whether models can solve problems, but how their reasoning unfolds. Unlike prior evaluations that measure only final-answer accuracy, PRISM-Bench introduces a diagnostic task: given a visual puzzle and a step-by-step chain-of-thought (CoT) containing exactly one error, models must identify the first incorrect step. This setting enables fine-grained assessment of logical consistency, error detection, and visual reasoning. The puzzles in PRISM-Bench require multi-step symbolic, geometric, and analogical reasoning, resisting shortcuts based on superficial pattern matching. Evaluations across state-of-the-art MLLMs reveal a persistent gap between fluent generation and faithful reasoning: models that produce plausible CoTs often fail to locate simple logical faults. By disentangling answer generation from reasoning verification, PRISM-Bench offers a sharper lens on multimodal reasoning competence and underscores the need for diagnostic evaluation protocols in the development of trustworthy MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
Qian, Yusu
Wan, Cheng
Jia, Chao
Yang, Yinfei
Zhao, Qingyu
Gan, Zhe
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) have achieved remarkable progress on vision-language tasks, yet their reasoning processes remain sometimes unreliable. We introduce PRISM-Bench, a benchmark of puzzle-based visual challenges designed to evaluate not only whether models can solve problems, but how their reasoning unfolds. Unlike prior evaluations that measure only final-answer accuracy, PRISM-Bench introduces a diagnostic task: given a visual puzzle and a step-by-step chain-of-thought (CoT) containing exactly one error, models must identify the first incorrect step. This setting enables fine-grained assessment of logical consistency, error detection, and visual reasoning. The puzzles in PRISM-Bench require multi-step symbolic, geometric, and analogical reasoning, resisting shortcuts based on superficial pattern matching. Evaluations across state-of-the-art MLLMs reveal a persistent gap between fluent generation and faithful reasoning: models that produce plausible CoTs often fail to locate simple logical faults. By disentangling answer generation from reasoning verification, PRISM-Bench offers a sharper lens on multimodal reasoning competence and underscores the need for diagnostic evaluation protocols in the development of trustworthy MLLMs.
title PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.23594