Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Qi, Zhuo, Le, Xin, Yi, Du, Ruoyi, Li, Zhen, Fu, Bin, Lu, Yiting, Yuan, Jiakang, Li, Xinyue, Liu, Dongyang, Zhu, Xiangyang, Zhang, Manyuan, Beddow, Will, Millon, Erwann, Perez, Victor, Wang, Wenhai, He, Conghui, Zhang, Bo, Liu, Xiaohong, Li, Hongsheng, Qiao, Yu, Xu, Chang, Gao, Peng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
by: Liu, Dongyang, et al.
Published: (2024)
by: Liu, Dongyang, et al.
Published: (2024)
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
by: Xin, Yi, et al.
Published: (2025)
by: Xin, Yi, et al.
Published: (2025)
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
by: Liu, Dongyang, et al.
Published: (2025)
by: Liu, Dongyang, et al.
Published: (2025)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
by: Lei, Jiayi, et al.
Published: (2025)
by: Lei, Jiayi, et al.
Published: (2025)
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
by: Gao, Peng, et al.
Published: (2024)
by: Gao, Peng, et al.
Published: (2024)
I-Max: Maximize the Resolution Potential of Pre-trained Rectified Flow Transformers with Projected Flow
by: Du, Ruoyi, et al.
Published: (2024)
by: Du, Ruoyi, et al.
Published: (2024)
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
by: Pu, Yuandong, et al.
Published: (2025)
by: Pu, Yuandong, et al.
Published: (2025)
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning
by: Zheng, Dian, et al.
Published: (2026)
by: Zheng, Dian, et al.
Published: (2026)
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
by: Xin, Yi, et al.
Published: (2025)
by: Xin, Yi, et al.
Published: (2025)
Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models
by: Wu, Xiaoshi, et al.
Published: (2024)
by: Wu, Xiaoshi, et al.
Published: (2024)
LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encoding
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
by: Zhao, Shitian, et al.
Published: (2025)
by: Zhao, Shitian, et al.
Published: (2025)
AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
by: Zheng, Dian, et al.
Published: (2025)
by: Zheng, Dian, et al.
Published: (2025)
Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
by: Liu, Dongyang, et al.
Published: (2025)
by: Liu, Dongyang, et al.
Published: (2025)
Lumina
Published: (2017)
Published: (2017)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
by: Chen, Zhekai, et al.
Published: (2026)
by: Chen, Zhekai, et al.
Published: (2026)
Embodied Image Compression
by: Li, Chunyi, et al.
Published: (2025)
by: Li, Chunyi, et al.
Published: (2025)
Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
by: Xin, Yi, et al.
Published: (2025)
by: Xin, Yi, et al.
Published: (2025)
Prompt Stealing Attacks Against Text-to-Image Generation Models
by: Shen, Xinyue, et al.
Published: (2023)
by: Shen, Xinyue, et al.
Published: (2023)
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
by: Image Team, et al.
Published: (2025)
by: Image Team, et al.
Published: (2025)
PICABench: How Far Are We from Physically Realistic Image Editing?
by: Pu, Yuandong, et al.
Published: (2025)
by: Pu, Yuandong, et al.
Published: (2025)
AdaTooler-V: Adaptive Tool-Use for Images and Videos
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
UMIT: Unifying Medical Imaging Tasks via Vision-Language Models
by: Yu, Haiyang, et al.
Published: (2025)
by: Yu, Haiyang, et al.
Published: (2025)
CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
by: Duan, Chengqi, et al.
Published: (2025)
by: Duan, Chengqi, et al.
Published: (2025)
VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning
by: Li, Zhong-Yu, et al.
Published: (2025)
by: Li, Zhong-Yu, et al.
Published: (2025)
Image Quality Assessment for Embodied AI
by: Li, Chunyi, et al.
Published: (2025)
by: Li, Chunyi, et al.
Published: (2025)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
by: Li, Kunchang, et al.
Published: (2022)
by: Li, Kunchang, et al.
Published: (2022)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning
by: Zhang, Chenhao, et al.
Published: (2026)
by: Zhang, Chenhao, et al.
Published: (2026)
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring
by: Peng, Xinge, et al.
Published: (2026)
by: Peng, Xinge, et al.
Published: (2026)
UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images
by: Qu, Yiting, et al.
Published: (2024)
by: Qu, Yiting, et al.
Published: (2024)
Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling
by: Shi, Xiaoyu, et al.
Published: (2024)
by: Shi, Xiaoyu, et al.
Published: (2024)
OmniCaptioner: One Captioner to Rule Them All
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
by: Zhao, Liangbing, et al.
Published: (2026)
by: Zhao, Liangbing, et al.
Published: (2026)
DiffuX2CT: Diffusion Learning to Reconstruct CT Images from Biplanar X-Rays
by: Liu, Xuhui, et al.
Published: (2024)
by: Liu, Xuhui, et al.
Published: (2024)
Simultaneous Determination of Young's Modulus and Density of Ultrathin Low‐k Films Using Surface Acoustic Waves
by: Li Zhang, et al.
Published: (2024)
by: Li Zhang, et al.
Published: (2024)
Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning
by: Shi, Fan, et al.
Published: (2025)
by: Shi, Fan, et al.
Published: (2025)
Similar Items
-
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
by: Liu, Dongyang, et al.
Published: (2024) -
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
by: Xin, Yi, et al.
Published: (2025) -
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
by: Zhuo, Le, et al.
Published: (2024) -
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
by: Liu, Dongyang, et al.
Published: (2025) -
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
by: Lei, Jiayi, et al.
Published: (2025)