ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918067845464064 |
|---|---|
| author | Chen, Junying Cai, Zhenyang Chen, Pengcheng Chen, Shunian Ji, Ke Wang, Xidong Yang, Yunjin Wang, Benyou |
| author_facet | Chen, Junying Cai, Zhenyang Chen, Pengcheng Chen, Shunian Ji, Ke Wang, Xidong Yang, Yunjin Wang, Benyou |
| contents | Recent advances in multimodal generative models have unlocked photorealistic, instruction-aligned image generation, yet leading systems like GPT-4o-Image remain proprietary and inaccessible. To democratize these capabilities, we present ShareGPT-4o-Image, the first dataset comprising 45K text-to-image and 46K text-and-image-to-image data, all synthesized using GPT-4o's image generation capabilities for distilling its advanced image generation abilities. Leveraging this dataset, we develop Janus-4o, a multimodal large language model capable of both text-to-image and text-and-image-to-image generation. Janus-4o not only significantly improves text-to-image generation over its predecessor, Janus-Pro, but also newly supports text-and-image-to-image generation. Notably, it achieves impressive performance in text-and-image-to-image generation from scratch, using only 91K synthetic samples and 6 hours of training on an 8 A800-GPU machine. We hope the release of ShareGPT-4o-Image and Janus-4o will foster open research in photorealistic, instruction-aligned image generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_18095 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Chen, Junying Cai, Zhenyang Chen, Pengcheng Chen, Shunian Ji, Ke Wang, Xidong Yang, Yunjin Wang, Benyou Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Recent advances in multimodal generative models have unlocked photorealistic, instruction-aligned image generation, yet leading systems like GPT-4o-Image remain proprietary and inaccessible. To democratize these capabilities, we present ShareGPT-4o-Image, the first dataset comprising 45K text-to-image and 46K text-and-image-to-image data, all synthesized using GPT-4o's image generation capabilities for distilling its advanced image generation abilities. Leveraging this dataset, we develop Janus-4o, a multimodal large language model capable of both text-to-image and text-and-image-to-image generation. Janus-4o not only significantly improves text-to-image generation over its predecessor, Janus-Pro, but also newly supports text-and-image-to-image generation. Notably, it achieves impressive performance in text-and-image-to-image generation from scratch, using only 91K synthetic samples and 6 hours of training on an 8 A800-GPU machine. We hope the release of ShareGPT-4o-Image and Janus-4o will foster open research in photorealistic, instruction-aligned image generation. |
| title | ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2506.18095 |