WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chow, Wei, Pan, Jiachun, Liang, Yongyuan, Zhou, Mingze, Song, Xue, Jia, Liyu, Zhang, Saining, Tang, Siliang, Li, Juncheng, Zhang, Fengda, Wu, Weijia, Zhang, Hanwang, Chua, Tat-Seng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908653563412480
author Chow, Wei
Pan, Jiachun
Liang, Yongyuan
Zhou, Mingze
Song, Xue
Jia, Liyu
Zhang, Saining
Tang, Siliang
Li, Juncheng
Zhang, Fengda
Wu, Weijia
Zhang, Hanwang
Chua, Tat-Seng
author_facet Chow, Wei
Pan, Jiachun
Liang, Yongyuan
Zhou, Mingze
Song, Xue
Jia, Liyu
Zhang, Saining
Tang, Siliang
Li, Juncheng
Zhang, Fengda
Wu, Weijia
Zhang, Hanwang
Chua, Tat-Seng
contents Recent advances in unified multimodal models (UMMs) have enabled impressive progress in visual comprehension and generation. However, existing datasets and benchmarks focus primarily on single-turn interactions, failing to capture the multi-turn, context-dependent nature of real-world image creation and editing. To address this gap, we present WEAVE, the first suite for in-context interleaved cross-modality comprehension and generation. Our suite consists of two complementary parts. WEAVE-100k is a large-scale dataset of 100K interleaved samples spanning over 370K dialogue turns and 500K images, covering comprehension, editing, and generation tasks that require reasoning over historical context. WEAVEBench is a human-annotated benchmark with 100 tasks based on 480 images, featuring a hybrid VLM judger evaluation framework based on both the reference image and the combination of the original image with editing instructions that assesses models' abilities in multi-turn generation, visual memory, and world-knowledge reasoning across diverse domains. Experiments demonstrate that training on WEAVE-100k enables vision comprehension, image editing, and comprehension-generation collaboration capabilities. Furthermore, it facilitates UMMs to develop emergent visual-memory capabilities, while extensive evaluations on WEAVEBench expose the persistent limitations and challenges of current approaches in multi-turn, context-aware image generation and editing. We believe WEAVE provides a view and foundation for studying in-context interleaved comprehension and generation for multi-modal community.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11434
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
Chow, Wei
Pan, Jiachun
Liang, Yongyuan
Zhou, Mingze
Song, Xue
Jia, Liyu
Zhang, Saining
Tang, Siliang
Li, Juncheng
Zhang, Fengda
Wu, Weijia
Zhang, Hanwang
Chua, Tat-Seng
Computer Vision and Pattern Recognition
Recent advances in unified multimodal models (UMMs) have enabled impressive progress in visual comprehension and generation. However, existing datasets and benchmarks focus primarily on single-turn interactions, failing to capture the multi-turn, context-dependent nature of real-world image creation and editing. To address this gap, we present WEAVE, the first suite for in-context interleaved cross-modality comprehension and generation. Our suite consists of two complementary parts. WEAVE-100k is a large-scale dataset of 100K interleaved samples spanning over 370K dialogue turns and 500K images, covering comprehension, editing, and generation tasks that require reasoning over historical context. WEAVEBench is a human-annotated benchmark with 100 tasks based on 480 images, featuring a hybrid VLM judger evaluation framework based on both the reference image and the combination of the original image with editing instructions that assesses models' abilities in multi-turn generation, visual memory, and world-knowledge reasoning across diverse domains. Experiments demonstrate that training on WEAVE-100k enables vision comprehension, image editing, and comprehension-generation collaboration capabilities. Furthermore, it facilitates UMMs to develop emergent visual-memory capabilities, while extensive evaluations on WEAVEBench expose the persistent limitations and challenges of current approaches in multi-turn, context-aware image generation and editing. We believe WEAVE provides a view and foundation for studying in-context interleaved comprehension and generation for multi-modal community.
title WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.11434