IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lei, Jiayi, Zhang, Renrui, Hu, Xiangfei, Lin, Weifeng, Li, Zhen, Sun, Wenjian, Du, Ruoyi, Zhuo, Le, Li, Zhongyu, Li, Xinyue, Zhao, Shitian, Guo, Ziyu, Lu, Yiting, Gao, Peng, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
by: Liu, Dongyang, et al.
Published: (2024)
by: Liu, Dongyang, et al.
Published: (2024)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
by: Lin, Weifeng, et al.
Published: (2025)
by: Lin, Weifeng, et al.
Published: (2025)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
by: Zhuo, Le, et al.
Published: (2025)
by: Zhuo, Le, et al.
Published: (2025)
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
by: Qin, Qi, et al.
Published: (2025)
by: Qin, Qi, et al.
Published: (2025)
I-Max: Maximize the Resolution Potential of Pre-trained Rectified Flow Transformers with Projected Flow
by: Du, Ruoyi, et al.
Published: (2024)
by: Du, Ruoyi, et al.
Published: (2024)
VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning
by: Li, Zhong-Yu, et al.
Published: (2025)
by: Li, Zhong-Yu, et al.
Published: (2025)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
by: Tong, Chengzhuo, et al.
Published: (2025)
by: Tong, Chengzhuo, et al.
Published: (2025)
LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
by: Zhao, Shitian, et al.
Published: (2025)
by: Zhao, Shitian, et al.
Published: (2025)
From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
by: Zhao, Liangbing, et al.
Published: (2026)
by: Zhao, Liangbing, et al.
Published: (2026)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
by: Chen, Xinyan, et al.
Published: (2025)
by: Chen, Xinyan, et al.
Published: (2025)
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
by: Huang, Victor Shea-Jay, et al.
Published: (2025)
by: Huang, Victor Shea-Jay, et al.
Published: (2025)
ECNet: Effective Controllable Text-to-Image Diffusion Models
by: Li, Sicheng, et al.
Published: (2024)
by: Li, Sicheng, et al.
Published: (2024)
IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
by: Zhang, Xikai, et al.
Published: (2025)
by: Zhang, Xikai, et al.
Published: (2025)
OmniCaptioner: One Captioner to Rule Them All
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
by: Zhang, Renrui, et al.
Published: (2023)
by: Zhang, Renrui, et al.
Published: (2023)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning
by: Zhang, Chenhao, et al.
Published: (2026)
by: Zhang, Chenhao, et al.
Published: (2026)
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models
by: Zhao, Shitian, et al.
Published: (2024)
by: Zhao, Shitian, et al.
Published: (2024)
Fine-grained Image Quality Assessment for Perceptual Image Restoration
by: Sheng, Xiangfei, et al.
Published: (2025)
by: Sheng, Xiangfei, et al.
Published: (2025)
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
by: Gao, Peng, et al.
Published: (2024)
by: Gao, Peng, et al.
Published: (2024)
Hybrid Agents for Image Restoration
by: Li, Bingchen, et al.
Published: (2025)
by: Li, Bingchen, et al.
Published: (2025)
LongT2IBench: A Benchmark for Evaluating Long Text-to-Image Generation with Graph-structured Annotations
by: Yang, Zhichao, et al.
Published: (2025)
by: Yang, Zhichao, et al.
Published: (2025)
DRNet: All-in-One Image Restoration via Prior-Guided Dynamic Reparameterization
by: Li, Ao, et al.
Published: (2026)
by: Li, Ao, et al.
Published: (2026)
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
by: Xu, Xinnuo, et al.
Published: (2025)
by: Xu, Xinnuo, et al.
Published: (2025)
3DRot: Rediscovering the Missing Primitive for RGB-Based 3D Augmentation
by: Yang, Shitian, et al.
Published: (2025)
by: Yang, Shitian, et al.
Published: (2025)
LossAgent: Towards Any Optimization Objectives for Image Processing with LLM Agents
by: Li, Bingchen, et al.
Published: (2024)
by: Li, Bingchen, et al.
Published: (2024)
Enhance the Image: Super Resolution using Artificial Intelligence in MRI
by: Li, Ziyu, et al.
Published: (2024)
by: Li, Ziyu, et al.
Published: (2024)
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
by: Xin, Yi, et al.
Published: (2025)
by: Xin, Yi, et al.
Published: (2025)
DDO-IN: Dual Domains Optimization for Implicit Neural Network to Eliminate Motion Artifact in Magnetic Resonance Imaging
by: Mai, Zhongyu, et al.
Published: (2025)
by: Mai, Zhongyu, et al.
Published: (2025)
Generating Multimodal Images with GAN: Integrating Text, Image, and Style
by: Tan, Chaoyi, et al.
Published: (2025)
by: Tan, Chaoyi, et al.
Published: (2025)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
by: Liu, Dongyang, et al.
Published: (2025)
by: Liu, Dongyang, et al.
Published: (2025)
CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
by: Wang, Guankun, et al.
Published: (2024)
by: Wang, Guankun, et al.
Published: (2024)
Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Similar Items
-
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
by: Liu, Dongyang, et al.
Published: (2024) -
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
by: Lin, Weifeng, et al.
Published: (2024) -
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
by: Jiang, Dongzhi, et al.
Published: (2025) -
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
by: Lin, Weifeng, et al.
Published: (2025) -
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)