Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Huanyu, Wu, Wenshan, Li, Chengzu, Shang, Ning, Xia, Yan, Huang, Yangyu, Zhang, Yifan, Dong, Li, Zhang, Zhang, Wang, Liang, Tan, Tieniu, Wei, Furu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912674109980672
author Zhang, Huanyu
Wu, Wenshan
Li, Chengzu
Shang, Ning
Xia, Yan
Huang, Yangyu
Zhang, Yifan
Dong, Li
Zhang, Zhang
Wang, Liang
Tan, Tieniu
Wei, Furu
author_facet Zhang, Huanyu
Wu, Wenshan
Li, Chengzu
Shang, Ning
Xia, Yan
Huang, Yangyu
Zhang, Yifan
Dong, Li
Zhang, Zhang
Wang, Liang
Tan, Tieniu
Wei, Furu
contents While Multimodal Large Language Models (MLLMs) excel at visual understanding, they often struggle in complex scenarios that require visual planning and imagination. Inspired by how humans use sketching as a form of visual thinking to develop and communicate ideas, we introduce Latent Sketchpad, a framework that equips MLLMs with an internal visual scratchpad. The internal visual representations of MLLMs have traditionally been confined to perceptual understanding. We repurpose them to support generative visual thought without compromising reasoning ability. Building on frontier MLLMs, our approach integrates visual generation directly into their native autoregressive reasoning process. It allows the model to interleave textual reasoning with the generation of visual latents. These latents guide the internal thought process and can be translated into sketch images for interpretability. To realize this, we introduce two components: a Context-Aware Vision Head autoregressively produces visual representations, and a pretrained Sketch Decoder renders these into human-interpretable images. We evaluate the framework on our new dataset MazePlanning. Experiments across various MLLMs show that Latent Sketchpad delivers comparable or even superior reasoning performance to their backbone. It further generalizes across distinct frontier MLLMs, including Gemma3 and Qwen2.5-VL. By extending model's textual reasoning to visual thinking, our framework opens new opportunities for richer human-computer interaction and broader applications. More details and resources are available on our project page: https://latent-sketchpad.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24514
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
Zhang, Huanyu
Wu, Wenshan
Li, Chengzu
Shang, Ning
Xia, Yan
Huang, Yangyu
Zhang, Yifan
Dong, Li
Zhang, Zhang
Wang, Liang
Tan, Tieniu
Wei, Furu
Computer Vision and Pattern Recognition
Computation and Language
While Multimodal Large Language Models (MLLMs) excel at visual understanding, they often struggle in complex scenarios that require visual planning and imagination. Inspired by how humans use sketching as a form of visual thinking to develop and communicate ideas, we introduce Latent Sketchpad, a framework that equips MLLMs with an internal visual scratchpad. The internal visual representations of MLLMs have traditionally been confined to perceptual understanding. We repurpose them to support generative visual thought without compromising reasoning ability. Building on frontier MLLMs, our approach integrates visual generation directly into their native autoregressive reasoning process. It allows the model to interleave textual reasoning with the generation of visual latents. These latents guide the internal thought process and can be translated into sketch images for interpretability. To realize this, we introduce two components: a Context-Aware Vision Head autoregressively produces visual representations, and a pretrained Sketch Decoder renders these into human-interpretable images. We evaluate the framework on our new dataset MazePlanning. Experiments across various MLLMs show that Latent Sketchpad delivers comparable or even superior reasoning performance to their backbone. It further generalizes across distinct frontier MLLMs, including Gemma3 and Qwen2.5-VL. By extending model's textual reasoning to visual thinking, our framework opens new opportunities for richer human-computer interaction and broader applications. More details and resources are available on our project page: https://latent-sketchpad.github.io/.
title Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2510.24514