ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Jiawei, Hao, Yunzhuo, Wang, Huichen Will, Li, Linjie, Shieh, Michael Qizhe, Choi, Yejin, Krishna, Ranjay, Cheng, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912933730058240
author Gu, Jiawei
Hao, Yunzhuo
Wang, Huichen Will
Li, Linjie
Shieh, Michael Qizhe
Choi, Yejin
Krishna, Ranjay
Cheng, Yu
author_facet Gu, Jiawei
Hao, Yunzhuo
Wang, Huichen Will
Li, Linjie
Shieh, Michael Qizhe
Choi, Yejin
Krishna, Ranjay
Cheng, Yu
contents Multimodal reasoning requires iterative coordination between language and vision, yet it remains unclear what constitutes a meaningful interleaved chain of thought. We posit that text and image thoughts should function as complementary rather than isomorphic modalities that mutually advance reasoning. Guided by this principle, we build ThinkMorph, a unified model fine-tuned on approximately 24K high-quality interleaved reasoning traces spanning tasks with varying visual engagement. ThinkMorph learns to generate progressive text-image reasoning steps that concretely manipulate visual content while maintaining coherent verbal logic. It delivers large gains on vision-centric benchmarks (averaging 34.7 percent over the base model) and generalizes to out-of-domain tasks, matching or surpassing larger and proprietary VLMs. Beyond performance, ThinkMorph exhibits emergent multimodal intelligence, including unseen visual manipulation skills, adaptive switching between reasoning modes, and better test-time scaling through diversified multimodal thoughts. These findings suggest promising directions for characterizing the emergent capabilities of unified models for multimodal reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
Gu, Jiawei
Hao, Yunzhuo
Wang, Huichen Will
Li, Linjie
Shieh, Michael Qizhe
Choi, Yejin
Krishna, Ranjay
Cheng, Yu
Computer Vision and Pattern Recognition
Multimodal reasoning requires iterative coordination between language and vision, yet it remains unclear what constitutes a meaningful interleaved chain of thought. We posit that text and image thoughts should function as complementary rather than isomorphic modalities that mutually advance reasoning. Guided by this principle, we build ThinkMorph, a unified model fine-tuned on approximately 24K high-quality interleaved reasoning traces spanning tasks with varying visual engagement. ThinkMorph learns to generate progressive text-image reasoning steps that concretely manipulate visual content while maintaining coherent verbal logic. It delivers large gains on vision-centric benchmarks (averaging 34.7 percent over the base model) and generalizes to out-of-domain tasks, matching or surpassing larger and proprietary VLMs. Beyond performance, ThinkMorph exhibits emergent multimodal intelligence, including unseen visual manipulation skills, adaptive switching between reasoning modes, and better test-time scaling through diversified multimodal thoughts. These findings suggest promising directions for characterizing the emergent capabilities of unified models for multimodal reasoning.
title ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.27492