OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zhihong, Bai, Xuehai, Shi, Yang, Fu, Chaoyou, Zhang, Huanyu, Wang, Haotian, Sun, Xiaoyan, Zhang, Zhang, Wang, Liang, Zhang, Yuanxing, Wan, Pengfei, Zhang, Yi-Fan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
by: Bai, Xuehai, et al.
Published: (2026)
by: Bai, Xuehai, et al.
Published: (2026)
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
by: Zhang, Huanyu, et al.
Published: (2026)
by: Zhang, Huanyu, et al.
Published: (2026)
Data Processing for the OpenGPT-X Model Family
by: Brandizzi, Nicolo', et al.
Published: (2024)
by: Brandizzi, Nicolo', et al.
Published: (2024)
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
by: Bai, Xuehai, et al.
Published: (2026)
by: Bai, Xuehai, et al.
Published: (2026)
BaseReward: A Strong Baseline for Multimodal Reward Model
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
by: Xie, Wulin, et al.
Published: (2025)
by: Xie, Wulin, et al.
Published: (2025)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
by: Wang, Qixun, et al.
Published: (2025)
by: Wang, Qixun, et al.
Published: (2025)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
by: Penke, Carolin, et al.
Published: (2025)
by: Penke, Carolin, et al.
Published: (2025)
Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
by: Zhang, Yi-Fan, et al.
Published: (2024)
by: Zhang, Yi-Fan, et al.
Published: (2024)
LogoRA: Local-Global Representation Alignment for Robust Time Series Classification
by: Zhang, Huanyu, et al.
Published: (2024)
by: Zhang, Huanyu, et al.
Published: (2024)
ReasonPix2Pix: Instruction Reasoning Dataset for Advanced Image Editing
by: Jin, Ying, et al.
Published: (2024)
by: Jin, Ying, et al.
Published: (2024)
LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
by: Wang, Weicheng, et al.
Published: (2026)
by: Wang, Weicheng, et al.
Published: (2026)
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
by: Wei, Xinyu, et al.
Published: (2025)
by: Wei, Xinyu, et al.
Published: (2025)
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
by: Zhang, Yi-Fan, et al.
Published: (2024)
by: Zhang, Yi-Fan, et al.
Published: (2024)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
Human Image Generation: A Comprehensive Survey
by: Jia, Zhen, et al.
Published: (2022)
by: Jia, Zhen, et al.
Published: (2022)
VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
by: Wang, Qunzhong, et al.
Published: (2025)
by: Wang, Qunzhong, et al.
Published: (2025)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
by: Zhou, Chenyu, et al.
Published: (2024)
by: Zhou, Chenyu, et al.
Published: (2024)
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
by: Tong, Chengzhuo, et al.
Published: (2026)
by: Tong, Chengzhuo, et al.
Published: (2026)
Thyme: Think Beyond Images
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
High-Fidelity GAN Inversion for Image Attribute Editing
by: Wang, Tengfei, et al.
Published: (2021)
by: Wang, Tengfei, et al.
Published: (2021)
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
by: Wang, Yuchi, et al.
Published: (2025)
by: Wang, Yuchi, et al.
Published: (2025)
Homoclinic points for convex billiards
by: Xia, Zhihong, et al.
Published: (2013)
by: Xia, Zhihong, et al.
Published: (2013)
MieDB-100k: A Comprehensive Dataset for Medical Image Editing
by: Lai, Yongfan, et al.
Published: (2026)
by: Lai, Yongfan, et al.
Published: (2026)
TimeRAF: Retrieval-Augmented Foundation model for Zero-shot Time Series Forecasting
by: Zhang, Huanyu, et al.
Published: (2024)
by: Zhang, Huanyu, et al.
Published: (2024)
NumeriKontrol: Adding Numeric Control to Diffusion Transformers for Instruction-based Image Editing
by: Xu, Zhenyu, et al.
Published: (2025)
by: Xu, Zhenyu, et al.
Published: (2025)
PSO‐SVM Machine Learning for Blasting Vibration Velocity Prediction in Open Pit Mines
by: Xuehai Chi, et al.
Published: (2026)
by: Xuehai Chi, et al.
Published: (2026)
Accelerated Schrödinger-Föllmer samplers
by: Lin, Haotian, et al.
Published: (2026)
by: Lin, Haotian, et al.
Published: (2026)
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
Image Matters: A New Dataset and Empirical Study for Multimodal Hyperbole Detection
by: Zhang, Huixuan, et al.
Published: (2023)
by: Zhang, Huixuan, et al.
Published: (2023)
An Attribute-Enriched Dataset and Auto-Annotated Pipeline for Open Detection
by: Qi, Pengfei, et al.
Published: (2024)
by: Qi, Pengfei, et al.
Published: (2024)
Recent Advances in Electrocatalytic Hydrogenation Reactions on Copper‐Based Catalysts
by: Min Zheng, et al.
Published: (2024)
by: Min Zheng, et al.
Published: (2024)
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
by: Wang, Yuhan, et al.
Published: (2025)
by: Wang, Yuhan, et al.
Published: (2025)
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
ControlEdit: A MultiModal Local Clothing Image Editing Method
by: Cheng, Di, et al.
Published: (2024)
by: Cheng, Di, et al.
Published: (2024)
Diffusion-Based Image Editing: An Unforeseen Adversary to Robust Invisible Watermarks
by: Fu, Wenkai, et al.
Published: (2025)
by: Fu, Wenkai, et al.
Published: (2025)
Denoising Designs-inherited Search Framework for Image Denoising
by: Zhang, Zheyu, et al.
Published: (2025)
by: Zhang, Zheyu, et al.
Published: (2025)
Routing to the Right Expertise: A Trustworthy Judge for Instruction-based Image Editing
by: Sun, Chenxi, et al.
Published: (2025)
by: Sun, Chenxi, et al.
Published: (2025)
RSEdit: Text-Guided Image Editing for Remote Sensing
by: Zhenyuan, Chen, et al.
Published: (2026)
by: Zhenyuan, Chen, et al.
Published: (2026)
Similar Items
-
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
by: Bai, Xuehai, et al.
Published: (2026) -
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
by: Zhang, Huanyu, et al.
Published: (2026) -
Data Processing for the OpenGPT-X Model Family
by: Brandizzi, Nicolo', et al.
Published: (2024) -
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
by: Bai, Xuehai, et al.
Published: (2026) -
BaseReward: A Strong Baseline for Multimodal Reward Model
by: Zhang, Yi-Fan, et al.
Published: (2025)