Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ye, Junyan, Jiang, Dongzhi, Wang, Zihao, Zhu, Leqi, Hu, Zhenghao, Huang, Zilong, He, Jun, Yan, Zhiyuan, Yu, Jinghua, Li, Hongsheng, He, Conghui, Li, Weijia
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916896647938048
author Ye, Junyan
Jiang, Dongzhi
Wang, Zihao
Zhu, Leqi
Hu, Zhenghao
Huang, Zilong
He, Jun
Yan, Zhiyuan
Yu, Jinghua
Li, Hongsheng
He, Conghui
Li, Weijia
author_facet Ye, Junyan
Jiang, Dongzhi
Wang, Zihao
Zhu, Leqi
Hu, Zhenghao
Huang, Zilong
He, Jun
Yan, Zhiyuan
Yu, Jinghua
Li, Hongsheng
He, Conghui
Li, Weijia
contents Recently, GPT-4o has garnered significant attention for its strong performance in image generation, yet open-source models still lag behind. Several studies have explored distilling image data from GPT-4o to enhance open-source models, achieving notable progress. However, a key question remains: given that real-world image datasets already constitute a natural source of high-quality data, why should we use GPT-4o-generated synthetic data? In this work, we identify two key advantages of synthetic images. First, they can complement rare scenarios in real-world datasets, such as surreal fantasy or multi-reference image generation, which frequently occur in user queries. Second, they provide clean and controllable supervision. Real-world data often contains complex background noise and inherent misalignment between text descriptions and image content, whereas synthetic images offer pure backgrounds and long-tailed supervision signals, facilitating more accurate text-to-image alignment. Building on these insights, we introduce Echo-4o-Image, a 180K-scale synthetic dataset generated by GPT-4o, harnessing the power of synthetic image data to address blind spots in real-world coverage. Using this dataset, we fine-tune the unified multimodal generation baseline Bagel to obtain Echo-4o. In addition, we propose two new evaluation benchmarks for a more accurate and challenging assessment of image generation capabilities: GenEval++, which increases instruction complexity to mitigate score saturation, and Imagine-Bench, which focuses on evaluating both the understanding and generation of imaginative content. Echo-4o demonstrates strong performance across standard benchmarks. Moreover, applying Echo-4o-Image to other foundation models (e.g., OmniGen2, BLIP3-o) yields consistent performance gains across multiple metrics, highlighting the datasets strong transferability.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09987
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
Ye, Junyan
Jiang, Dongzhi
Wang, Zihao
Zhu, Leqi
Hu, Zhenghao
Huang, Zilong
He, Jun
Yan, Zhiyuan
Yu, Jinghua
Li, Hongsheng
He, Conghui
Li, Weijia
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Recently, GPT-4o has garnered significant attention for its strong performance in image generation, yet open-source models still lag behind. Several studies have explored distilling image data from GPT-4o to enhance open-source models, achieving notable progress. However, a key question remains: given that real-world image datasets already constitute a natural source of high-quality data, why should we use GPT-4o-generated synthetic data? In this work, we identify two key advantages of synthetic images. First, they can complement rare scenarios in real-world datasets, such as surreal fantasy or multi-reference image generation, which frequently occur in user queries. Second, they provide clean and controllable supervision. Real-world data often contains complex background noise and inherent misalignment between text descriptions and image content, whereas synthetic images offer pure backgrounds and long-tailed supervision signals, facilitating more accurate text-to-image alignment. Building on these insights, we introduce Echo-4o-Image, a 180K-scale synthetic dataset generated by GPT-4o, harnessing the power of synthetic image data to address blind spots in real-world coverage. Using this dataset, we fine-tune the unified multimodal generation baseline Bagel to obtain Echo-4o. In addition, we propose two new evaluation benchmarks for a more accurate and challenging assessment of image generation capabilities: GenEval++, which increases instruction complexity to mitigate score saturation, and Imagine-Bench, which focuses on evaluating both the understanding and generation of imaginative content. Echo-4o demonstrates strong performance across standard benchmarks. Moreover, applying Echo-4o-Image to other foundation models (e.g., OmniGen2, BLIP3-o) yields consistent performance gains across multiple metrics, highlighting the datasets strong transferability.
title Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.09987