ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junying, Cai, Zhenyang, Chen, Pengcheng, Chen, Shunian, Ji, Ke, Wang, Xidong, Yang, Yunjin, Wang, Benyou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918067845464064
author Chen, Junying
Cai, Zhenyang
Chen, Pengcheng
Chen, Shunian
Ji, Ke
Wang, Xidong
Yang, Yunjin
Wang, Benyou
author_facet Chen, Junying
Cai, Zhenyang
Chen, Pengcheng
Chen, Shunian
Ji, Ke
Wang, Xidong
Yang, Yunjin
Wang, Benyou
contents Recent advances in multimodal generative models have unlocked photorealistic, instruction-aligned image generation, yet leading systems like GPT-4o-Image remain proprietary and inaccessible. To democratize these capabilities, we present ShareGPT-4o-Image, the first dataset comprising 45K text-to-image and 46K text-and-image-to-image data, all synthesized using GPT-4o's image generation capabilities for distilling its advanced image generation abilities. Leveraging this dataset, we develop Janus-4o, a multimodal large language model capable of both text-to-image and text-and-image-to-image generation. Janus-4o not only significantly improves text-to-image generation over its predecessor, Janus-Pro, but also newly supports text-and-image-to-image generation. Notably, it achieves impressive performance in text-and-image-to-image generation from scratch, using only 91K synthetic samples and 6 hours of training on an 8 A800-GPU machine. We hope the release of ShareGPT-4o-Image and Janus-4o will foster open research in photorealistic, instruction-aligned image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18095
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Chen, Junying
Cai, Zhenyang
Chen, Pengcheng
Chen, Shunian
Ji, Ke
Wang, Xidong
Yang, Yunjin
Wang, Benyou
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recent advances in multimodal generative models have unlocked photorealistic, instruction-aligned image generation, yet leading systems like GPT-4o-Image remain proprietary and inaccessible. To democratize these capabilities, we present ShareGPT-4o-Image, the first dataset comprising 45K text-to-image and 46K text-and-image-to-image data, all synthesized using GPT-4o's image generation capabilities for distilling its advanced image generation abilities. Leveraging this dataset, we develop Janus-4o, a multimodal large language model capable of both text-to-image and text-and-image-to-image generation. Janus-4o not only significantly improves text-to-image generation over its predecessor, Janus-Pro, but also newly supports text-and-image-to-image generation. Notably, it achieves impressive performance in text-and-image-to-image generation from scratch, using only 91K synthetic samples and 6 hours of training on an 8 A800-GPU machine. We hope the release of ShareGPT-4o-Image and Janus-4o will foster open research in photorealistic, instruction-aligned image generation.
title ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.18095