OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Zhihong, Bai, Xuehai, Shi, Yang, Fu, Chaoyou, Zhang, Huanyu, Wang, Haotian, Sun, Xiaoyan, Zhang, Zhang, Wang, Liang, Zhang, Yuanxing, Wan, Pengfei, Zhang, Yi-Fan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912615131774976
author Chen, Zhihong
Bai, Xuehai
Shi, Yang
Fu, Chaoyou
Zhang, Huanyu
Wang, Haotian
Sun, Xiaoyan
Zhang, Zhang
Wang, Liang
Zhang, Yuanxing
Wan, Pengfei
Zhang, Yi-Fan
author_facet Chen, Zhihong
Bai, Xuehai
Shi, Yang
Fu, Chaoyou
Zhang, Huanyu
Wang, Haotian
Sun, Xiaoyan
Zhang, Zhang
Wang, Liang
Zhang, Yuanxing
Wan, Pengfei
Zhang, Yi-Fan
contents The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and simple object manipulation, they often lack the systematic structure and challenging scenarios required for real-world applications. To address this bottleneck, we introduce OpenGPT-4o-Image, a large-scale dataset constructed using a novel methodology that combines hierarchical task taxonomy with automated data generation. Our taxonomy not only includes fundamental capabilities such as text rendering and style control but also introduces highly practical yet challenging categories like scientific imagery for chemistry illustrations and complex instruction editing requiring simultaneous execution of multiple operations. Through an automated pipeline leveraging structured resource pools and GPT-4o, we generate 80k high-quality instruction-image pairs with controlled diversity, covering 11 major domains and 51 subtasks. Extensive experiments show that fine-tuning leading models on our dataset achieves significant performance gains across multiple benchmarks, with improvements of up to 18\% on editing tasks (UniWorld-V1 on ImgEdit-Bench) and 13% on generation tasks (Harmon on GenEval). Our work demonstrates that systematic data construction is key to advancing multimodal AI capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24900
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
Chen, Zhihong
Bai, Xuehai
Shi, Yang
Fu, Chaoyou
Zhang, Huanyu
Wang, Haotian
Sun, Xiaoyan
Zhang, Zhang
Wang, Liang
Zhang, Yuanxing
Wan, Pengfei
Zhang, Yi-Fan
Computer Vision and Pattern Recognition
Artificial Intelligence
The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and simple object manipulation, they often lack the systematic structure and challenging scenarios required for real-world applications. To address this bottleneck, we introduce OpenGPT-4o-Image, a large-scale dataset constructed using a novel methodology that combines hierarchical task taxonomy with automated data generation. Our taxonomy not only includes fundamental capabilities such as text rendering and style control but also introduces highly practical yet challenging categories like scientific imagery for chemistry illustrations and complex instruction editing requiring simultaneous execution of multiple operations. Through an automated pipeline leveraging structured resource pools and GPT-4o, we generate 80k high-quality instruction-image pairs with controlled diversity, covering 11 major domains and 51 subtasks. Extensive experiments show that fine-tuning leading models on our dataset achieves significant performance gains across multiple benchmarks, with improvements of up to 18\% on editing tasks (UniWorld-V1 on ImgEdit-Bench) and 13% on generation tasks (Harmon on GenEval). Our work demonstrates that systematic data construction is key to advancing multimodal AI capabilities.
title OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.24900