GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yuhan, Yang, Siwei, Zhao, Bingchen, Zhang, Letian, Liu, Qing, Zhou, Yuyin, Xie, Cihang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
by: Yang, Siwei, et al.
Published: (2025)
by: Yang, Siwei, et al.
Published: (2025)
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
by: Hui, Mude, et al.
Published: (2024)
by: Hui, Mude, et al.
Published: (2024)
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
by: Yang, Jinrui, et al.
Published: (2026)
by: Yang, Jinrui, et al.
Published: (2026)
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
by: Liu, Yanqing, et al.
Published: (2025)
by: Liu, Yanqing, et al.
Published: (2025)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
3D-TransUNet for Brain Metastases Segmentation in the BraTS2023 Challenge
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
Scaling White-Box Transformers for Vision
by: Yang, Jinrui, et al.
Published: (2024)
by: Yang, Jinrui, et al.
Published: (2024)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
by: Gao, Yipeng, et al.
Published: (2023)
by: Gao, Yipeng, et al.
Published: (2023)
Generative Image Layer Decomposition with Visual Effects
by: Yang, Jinrui, et al.
Published: (2024)
by: Yang, Jinrui, et al.
Published: (2024)
Rejuvenating image-GPT as Strong Visual Representation Learners
by: Ren, Sucheng, et al.
Published: (2023)
by: Ren, Sucheng, et al.
Published: (2023)
Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation
by: Zhou, Yufan, et al.
Published: (2024)
by: Zhou, Yufan, et al.
Published: (2024)
What If We Recaption Billions of Web Images with LLaMA-3?
by: Li, Xianhang, et al.
Published: (2024)
by: Li, Xianhang, et al.
Published: (2024)
Oasis: One Image is All You Need for Multimodal Instruction Data Synthesis
by: Zhang, Letian, et al.
Published: (2025)
by: Zhang, Letian, et al.
Published: (2025)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
by: Liu, Yanqing, et al.
Published: (2024)
by: Liu, Yanqing, et al.
Published: (2024)
MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine
by: Xie, Yunfei, et al.
Published: (2024)
by: Xie, Yunfei, et al.
Published: (2024)
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
by: Mao, Jiawei, et al.
Published: (2026)
by: Mao, Jiawei, et al.
Published: (2026)
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset
by: Liu, Jiazhen, et al.
Published: (2024)
by: Liu, Jiazhen, et al.
Published: (2024)
FingerVeinSyn-5M: A Million-Scale Dataset and Benchmark for Finger Vein Recognition
by: Wang, Yinfan, et al.
Published: (2025)
by: Wang, Yinfan, et al.
Published: (2025)
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
by: Peng, Cihang, et al.
Published: (2025)
by: Peng, Cihang, et al.
Published: (2025)
A Preliminary Study on GPT-Image Generation Model for Image Restoration
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
An Empirical Study of GPT-4o Image Generation Capabilities
by: Chen, Sixiang, et al.
Published: (2025)
by: Chen, Sixiang, et al.
Published: (2025)
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
by: Xu, Linrui, et al.
Published: (2024)
by: Xu, Linrui, et al.
Published: (2024)
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
by: Zhao, Bingchen, et al.
Published: (2024)
by: Zhao, Bingchen, et al.
Published: (2024)
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
by: Yan, Zhiyuan, et al.
Published: (2025)
by: Yan, Zhiyuan, et al.
Published: (2025)
GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment
by: Zewde, Kidus, et al.
Published: (2026)
by: Zewde, Kidus, et al.
Published: (2026)
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
by: Chen, Zhihong, et al.
Published: (2025)
by: Chen, Zhihong, et al.
Published: (2025)
LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images
by: Flora, James, et al.
Published: (2026)
by: Flora, James, et al.
Published: (2026)
RadGPT: Constructing 3D Image-Text Tumor Datasets
by: Bassi, Pedro R. A. S., et al.
Published: (2025)
by: Bassi, Pedro R. A. S., et al.
Published: (2025)
CoDA: From Text-to-Image Diffusion Models to Training-Free Dataset Distillation
by: Zhou, Letian, et al.
Published: (2025)
by: Zhou, Letian, et al.
Published: (2025)
Preliminary Explorations with GPT-4o(mni) Native Image Generation
by: Cao, Pu, et al.
Published: (2025)
by: Cao, Pu, et al.
Published: (2025)
ARFlow: Autoregressive Flow with Hybrid Linear Attention
by: Hui, Mude, et al.
Published: (2025)
by: Hui, Mude, et al.
Published: (2025)
Revisiting Adversarial Training at Scale
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
SIGMAN:Scaling 3D Human Gaussian Generation with Millions of Assets
by: Yang, Yuhang, et al.
Published: (2025)
by: Yang, Yuhang, et al.
Published: (2025)
Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies
by: Li, Zichao, et al.
Published: (2024)
by: Li, Zichao, et al.
Published: (2024)
MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
by: Song, Kunpeng, et al.
Published: (2024)
by: Song, Kunpeng, et al.
Published: (2024)
MMO-IG: Multi-Class and Multi-Scale Object Image Generation for Remote Sensing
by: Yang, Chuang, et al.
Published: (2024)
by: Yang, Chuang, et al.
Published: (2024)
AnimeDL-2M: Million-Scale AI-Generated Anime Image Detection and Localization in Diffusion Era
by: Zhu, Chenyang, et al.
Published: (2025)
by: Zhu, Chenyang, et al.
Published: (2025)
Similar Items
-
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
by: Yang, Siwei, et al.
Published: (2025) -
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
by: Hui, Mude, et al.
Published: (2024) -
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
by: Yang, Jinrui, et al.
Published: (2026) -
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
by: Liu, Yanqing, et al.
Published: (2025) -
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
by: Wang, Feng, et al.
Published: (2025)