UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Pengyu, Zhou, Shaojun, Tan, Chenkun, Wang, Xinghao, Huang, Wei, Ye, Zhen, Li, Zhaowei, Jiang, Botian, Zhang, Dong, Qiu, Xipeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909795231989760
author Wang, Pengyu
Zhou, Shaojun
Tan, Chenkun
Wang, Xinghao
Huang, Wei
Ye, Zhen
Li, Zhaowei
Jiang, Botian
Zhang, Dong
Qiu, Xipeng
author_facet Wang, Pengyu
Zhou, Shaojun
Tan, Chenkun
Wang, Xinghao
Huang, Wei
Ye, Zhen
Li, Zhaowei
Jiang, Botian
Zhang, Dong
Qiu, Xipeng
contents Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image synthesis. However, progress in unified VLLMs remains constrained by the lack of datasets that fully exploit the synergistic potential between these two core abilities. Existing datasets typically address understanding and generation in isolation, thereby limiting the performance of unified VLLMs. To bridge this critical gap, we introduce a novel dataset construction framework, UnifiedVisual, and present UnifiedVisual-240K, a high-quality dataset meticulously designed to facilitate mutual enhancement between multimodal understanding and generation. UnifiedVisual-240K seamlessly integrates diverse visual and textual inputs and outputs, enabling comprehensive cross-modal reasoning and precise text-to-image alignment. Our dataset encompasses a wide spectrum of tasks and data sources, ensuring rich diversity and addressing key shortcomings of prior resources. Extensive experiments demonstrate that models trained on UnifiedVisual-240K consistently achieve strong performance across a wide range of tasks. Notably, these models exhibit significant mutual reinforcement between multimodal understanding and generation, further validating the effectiveness of our framework and dataset. We believe UnifiedVisual represents a new growth point for advancing unified VLLMs and unlocking their full potential. Our code and datasets is available at https://github.com/fnlp-vision/UnifiedVisual.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14738
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
Wang, Pengyu
Zhou, Shaojun
Tan, Chenkun
Wang, Xinghao
Huang, Wei
Ye, Zhen
Li, Zhaowei
Jiang, Botian
Zhang, Dong
Qiu, Xipeng
Computation and Language
Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image synthesis. However, progress in unified VLLMs remains constrained by the lack of datasets that fully exploit the synergistic potential between these two core abilities. Existing datasets typically address understanding and generation in isolation, thereby limiting the performance of unified VLLMs. To bridge this critical gap, we introduce a novel dataset construction framework, UnifiedVisual, and present UnifiedVisual-240K, a high-quality dataset meticulously designed to facilitate mutual enhancement between multimodal understanding and generation. UnifiedVisual-240K seamlessly integrates diverse visual and textual inputs and outputs, enabling comprehensive cross-modal reasoning and precise text-to-image alignment. Our dataset encompasses a wide spectrum of tasks and data sources, ensuring rich diversity and addressing key shortcomings of prior resources. Extensive experiments demonstrate that models trained on UnifiedVisual-240K consistently achieve strong performance across a wide range of tasks. Notably, these models exhibit significant mutual reinforcement between multimodal understanding and generation, further validating the effectiveness of our framework and dataset. We believe UnifiedVisual represents a new growth point for advancing unified VLLMs and unlocking their full potential. Our code and datasets is available at https://github.com/fnlp-vision/UnifiedVisual.
title UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
topic Computation and Language
url https://arxiv.org/abs/2509.14738