Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Yifan, Hu, Liang, Sun, Haofeng, Wang, Peiyu, Wei, Yichen, Yin, Shukang, Pei, Jiangbo, Shen, Wei, Xia, Peng, Peng, Yi, Xie, Tianyidan, Li, Eric, Liu, Yang, Song, Xuchen, Zhou, Yahui
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915658831233024
author Zhang, Yifan
Hu, Liang
Sun, Haofeng
Wang, Peiyu
Wei, Yichen
Yin, Shukang
Pei, Jiangbo
Shen, Wei
Xia, Peng
Peng, Yi
Xie, Tianyidan
Li, Eric
Liu, Yang
Song, Xuchen
Zhou, Yahui
author_facet Zhang, Yifan
Hu, Liang
Sun, Haofeng
Wang, Peiyu
Wei, Yichen
Yin, Shukang
Pei, Jiangbo
Shen, Wei
Xia, Peng
Peng, Yi
Xie, Tianyidan
Li, Eric
Liu, Yang
Song, Xuchen
Zhou, Yahui
contents Despite recent progress in multimodal agentic systems, existing approaches often treat image manipulation and web search as disjoint capabilities, rely heavily on costly reinforcement learning, and lack planning grounded in real tool-execution traces. To address these limitations, we present Skywork-R1V4, a 30B (A3B) parameter multimodal agentic model that unifies multimodal planning, active image manipulation ("thinking with images"), deep multimodal search, and, most critically, interleaved reasoning that dynamically alternates between visual operations and external knowledge retrieval. Trained solely via supervised fine-tuning on fewer than 30,000 high-quality, planning-execution-consistent trajectories and validated through stepwise consistency filtering, Skywork-R1V4 achieves state-of-the-art results across perception and multimodal search benchmarks: it scores 66.1 on MMSearch and 67.2 on FVQA, surpassing Gemini 2.5 Flash on all 11 metrics. Skywork-R1V4 exhibits emergent long-horizon reasoning at inference time, successfully orchestrating more than 10 tool calls to solve complex, multi-step tasks. Our results demonstrate that sophisticated agentic multimodal intelligence can be achieved through carefully curated supervised learning alone, without any reliance on reinforcement learning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02395
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
Zhang, Yifan
Hu, Liang
Sun, Haofeng
Wang, Peiyu
Wei, Yichen
Yin, Shukang
Pei, Jiangbo
Shen, Wei
Xia, Peng
Peng, Yi
Xie, Tianyidan
Li, Eric
Liu, Yang
Song, Xuchen
Zhou, Yahui
Computer Vision and Pattern Recognition
Despite recent progress in multimodal agentic systems, existing approaches often treat image manipulation and web search as disjoint capabilities, rely heavily on costly reinforcement learning, and lack planning grounded in real tool-execution traces. To address these limitations, we present Skywork-R1V4, a 30B (A3B) parameter multimodal agentic model that unifies multimodal planning, active image manipulation ("thinking with images"), deep multimodal search, and, most critically, interleaved reasoning that dynamically alternates between visual operations and external knowledge retrieval. Trained solely via supervised fine-tuning on fewer than 30,000 high-quality, planning-execution-consistent trajectories and validated through stepwise consistency filtering, Skywork-R1V4 achieves state-of-the-art results across perception and multimodal search benchmarks: it scores 66.1 on MMSearch and 67.2 on FVQA, surpassing Gemini 2.5 Flash on all 11 metrics. Skywork-R1V4 exhibits emergent long-horizon reasoning at inference time, successfully orchestrating more than 10 tool calls to solve complex, multi-step tasks. Our results demonstrate that sophisticated agentic multimodal intelligence can be achieved through carefully curated supervised learning alone, without any reliance on reinforcement learning.
title Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.02395