Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Acuna, David, Lu, Ximing, Jung, Jaehun, Kim, Hyunwoo, Kar, Amlan, Fidler, Sanja, Choi, Yejin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908402194579456
author Acuna, David
Lu, Ximing
Jung, Jaehun
Kim, Hyunwoo
Kar, Amlan
Fidler, Sanja
Choi, Yejin
author_facet Acuna, David
Lu, Ximing
Jung, Jaehun
Kim, Hyunwoo
Kar, Amlan
Fidler, Sanja
Choi, Yejin
contents Recent research in vision-language models (VLMs) has centered around the possibility of equipping them with implicit long-form chain-of-thought reasoning -- akin to the success observed in language models -- via distillation and reinforcement learning. But what about the non-reasoning models already trained and deployed across the internet? Should we simply abandon them, or is there hope for a search mechanism that can elicit hidden knowledge and induce long reasoning traces -- without any additional training or supervision? In this paper, we explore this possibility using a Monte Carlo Tree Search (MCTS)-inspired algorithm, which injects subquestion-subanswer pairs into the model's output stream. We show that framing reasoning as a search process -- where subquestions act as latent decisions within a broader inference trajectory -- helps the model "connect the dots" between fragmented knowledge and produce extended reasoning traces in non-reasoning models. We evaluate our method across three benchmarks and observe consistent improvements. Notably, our approach yields a 2% overall improvement on MMMU-PRO, including a significant 9% gain in Liberal Arts.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08927
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
Acuna, David
Lu, Ximing
Jung, Jaehun
Kim, Hyunwoo
Kar, Amlan
Fidler, Sanja
Choi, Yejin
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Recent research in vision-language models (VLMs) has centered around the possibility of equipping them with implicit long-form chain-of-thought reasoning -- akin to the success observed in language models -- via distillation and reinforcement learning. But what about the non-reasoning models already trained and deployed across the internet? Should we simply abandon them, or is there hope for a search mechanism that can elicit hidden knowledge and induce long reasoning traces -- without any additional training or supervision? In this paper, we explore this possibility using a Monte Carlo Tree Search (MCTS)-inspired algorithm, which injects subquestion-subanswer pairs into the model's output stream. We show that framing reasoning as a search process -- where subquestions act as latent decisions within a broader inference trajectory -- helps the model "connect the dots" between fragmented knowledge and produce extended reasoning traces in non-reasoning models. We evaluate our method across three benchmarks and observe consistent improvements. Notably, our approach yields a 2% overall improvement on MMMU-PRO, including a significant 9% gain in Liberal Arts.
title Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.08927