Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Dongyang, Zhao, Shitian, Zhuo, Le, Lin, Weifeng, Xin, Yi, Li, Xinyue, Qin, Qi, Qiao, Yu, Li, Hongsheng, Gao, Peng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
por: Xin, Yi, et al.
Publicado: (2025)
por: Xin, Yi, et al.
Publicado: (2025)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
por: Liu, Dongyang, et al.
Publicado: (2025)
por: Liu, Dongyang, et al.
Publicado: (2025)
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
por: Qin, Qi, et al.
Publicado: (2025)
por: Qin, Qi, et al.
Publicado: (2025)
LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
por: Zhao, Shitian, et al.
Publicado: (2025)
por: Zhao, Shitian, et al.
Publicado: (2025)
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
por: Zhuo, Le, et al.
Publicado: (2024)
por: Zhuo, Le, et al.
Publicado: (2024)
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
por: Pu, Yuandong, et al.
Publicado: (2025)
por: Pu, Yuandong, et al.
Publicado: (2025)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
por: Lei, Jiayi, et al.
Publicado: (2025)
por: Lei, Jiayi, et al.
Publicado: (2025)
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
por: Gao, Peng, et al.
Publicado: (2024)
por: Gao, Peng, et al.
Publicado: (2024)
I-Max: Maximize the Resolution Potential of Pre-trained Rectified Flow Transformers with Projected Flow
por: Du, Ruoyi, et al.
Publicado: (2024)
por: Du, Ruoyi, et al.
Publicado: (2024)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
por: Lin, Weifeng, et al.
Publicado: (2024)
por: Lin, Weifeng, et al.
Publicado: (2024)
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
por: Xin, Yi, et al.
Publicado: (2025)
por: Xin, Yi, et al.
Publicado: (2025)
Lumina
Publicado: (2017)
Publicado: (2017)
Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
por: Xin, Yi, et al.
Publicado: (2025)
por: Xin, Yi, et al.
Publicado: (2025)
FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation
por: He, Xuehai, et al.
Publicado: (2024)
por: He, Xuehai, et al.
Publicado: (2024)
OmniCaptioner: One Captioner to Rule Them All
por: Lu, Yiting, et al.
Publicado: (2025)
por: Lu, Yiting, et al.
Publicado: (2025)
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
por: Ye, Junyan, et al.
Publicado: (2025)
por: Ye, Junyan, et al.
Publicado: (2025)
Emu: Generative Pretraining in Multimodality
por: Sun, Quan, et al.
Publicado: (2023)
por: Sun, Quan, et al.
Publicado: (2023)
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
por: Zhuo, Le, et al.
Publicado: (2025)
por: Zhuo, Le, et al.
Publicado: (2025)
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
por: Huang, Victor Shea-Jay, et al.
Publicado: (2025)
por: Huang, Victor Shea-Jay, et al.
Publicado: (2025)
M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
por: Luo, Mingshuang, et al.
Publicado: (2024)
por: Luo, Mingshuang, et al.
Publicado: (2024)
Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
por: Liang, Xinyue, et al.
Publicado: (2025)
por: Liang, Xinyue, et al.
Publicado: (2025)
POLAR: A Portrait OLAT Dataset and Generative Framework for Illumination-Aware Face Modeling
por: Chen, Zhuo, et al.
Publicado: (2025)
por: Chen, Zhuo, et al.
Publicado: (2025)
Generative Head-Mounted Camera Captures for Photorealistic Avatars
por: Bai, Shaojie, et al.
Publicado: (2025)
por: Bai, Shaojie, et al.
Publicado: (2025)
Is ChatGPT Involved in Texts? Measure the Polish Ratio to Detect ChatGPT-Generated Text
por: Yang, Lingyi, et al.
Publicado: (2023)
por: Yang, Lingyi, et al.
Publicado: (2023)
Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text Transformers
por: Gong, Linyuan, et al.
Publicado: (2023)
por: Gong, Linyuan, et al.
Publicado: (2023)
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
por: Liu, Dongyang, et al.
Publicado: (2024)
por: Liu, Dongyang, et al.
Publicado: (2024)
PE: A Poincare Explanation Method for Fast Text Hierarchy Generation
por: Chen, Qian, et al.
Publicado: (2024)
por: Chen, Qian, et al.
Publicado: (2024)
Exploring Flexible Scenario Generation in Godot Simulator
por: Peraltai, Daniel, et al.
Publicado: (2024)
por: Peraltai, Daniel, et al.
Publicado: (2024)
PagPassGPT: Pattern Guided Password Guessing via Generative Pretrained Transformer
por: Su, Xingyu, et al.
Publicado: (2024)
por: Su, Xingyu, et al.
Publicado: (2024)
NetGPT: Generative Pretrained Transformer for Network Traffic
por: Meng, Xuying, et al.
Publicado: (2023)
por: Meng, Xuying, et al.
Publicado: (2023)
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation
por: Izzati, Fathinah, et al.
Publicado: (2025)
por: Izzati, Fathinah, et al.
Publicado: (2025)
Semantic–Electromagnetic Inversion With Pretrained Multimodal Generative Model
por: Yanjin Chen, et al.
Publicado: (2024)
por: Yanjin Chen, et al.
Publicado: (2024)
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
por: Pan, Kaihang, et al.
Publicado: (2025)
por: Pan, Kaihang, et al.
Publicado: (2025)
Uncovering Regional Defaults from Photorealistic Forests in Text-to-Image Generation with DALL-E 2
por: Liu, Zilong, et al.
Publicado: (2024)
por: Liu, Zilong, et al.
Publicado: (2024)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
por: Xue, Jinlong, et al.
Publicado: (2024)
por: Xue, Jinlong, et al.
Publicado: (2024)
VEnhancer: Generative Space-Time Enhancement for Video Generation
por: He, Jingwen, et al.
Publicado: (2024)
por: He, Jingwen, et al.
Publicado: (2024)
Mechanical‐Thermal Co‐Design of Flexible Thermoelectric Devices with Solid‐Liquid Electrodes for Enhanced Stretchability and Power Generation
por: Keke Chen, et al.
Publicado: (2025)
por: Keke Chen, et al.
Publicado: (2025)
Why Compress What You Can Generate? When GPT-4o Generation Ushers in Image Compression Fields
por: Gao, Yixin, et al.
Publicado: (2025)
por: Gao, Yixin, et al.
Publicado: (2025)
Boundless: Generating Photorealistic Synthetic Data for Object Detection in Urban Streetscapes
por: Turkcan, Mehmet Kerem, et al.
Publicado: (2024)
por: Turkcan, Mehmet Kerem, et al.
Publicado: (2024)
CSI-GPT: Integrating Generative Pre-Trained Transformer with Federated-Tuning to Acquire Downlink Massive MIMO Channels
por: Zeng, Ye, et al.
Publicado: (2024)
por: Zeng, Ye, et al.
Publicado: (2024)
Ejemplares similares
-
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
por: Xin, Yi, et al.
Publicado: (2025) -
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
por: Liu, Dongyang, et al.
Publicado: (2025) -
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
por: Qin, Qi, et al.
Publicado: (2025) -
LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
por: Zhao, Shitian, et al.
Publicado: (2025) -
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
por: Zhuo, Le, et al.
Publicado: (2024)