Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhuo, Le, Du, Ruoyi, Xiao, Han, Li, Yangguang, Liu, Dongyang, Huang, Rongjie, Liu, Wenze, Zhao, Lirui, Wang, Fu-Yun, Ma, Zhanyu, Luo, Xu, Wang, Zehan, Zhang, Kaipeng, Zhu, Xiangyang, Liu, Si, Yue, Xiangyu, Liu, Dingning, Ouyang, Wanli, Liu, Ziwei, Qiao, Yu, Li, Hongsheng, Gao, Peng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
von: Liu, Dongyang, et al.
Veröffentlicht: (2025)
von: Liu, Dongyang, et al.
Veröffentlicht: (2025)
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
von: Qin, Qi, et al.
Veröffentlicht: (2025)
von: Qin, Qi, et al.
Veröffentlicht: (2025)
Lumina
Veröffentlicht: (2017)
Veröffentlicht: (2017)
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
von: Liu, Dongyang, et al.
Veröffentlicht: (2024)
von: Liu, Dongyang, et al.
Veröffentlicht: (2024)
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
von: Gao, Peng, et al.
Veröffentlicht: (2024)
von: Gao, Peng, et al.
Veröffentlicht: (2024)
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
von: Xin, Yi, et al.
Veröffentlicht: (2025)
von: Xin, Yi, et al.
Veröffentlicht: (2025)
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
von: Xin, Yi, et al.
Veröffentlicht: (2025)
von: Xin, Yi, et al.
Veröffentlicht: (2025)
MeshCraft: Exploring Efficient and Controllable Mesh Generation with Flow-based DiTs
von: He, Xianglong, et al.
Veröffentlicht: (2025)
von: He, Xianglong, et al.
Veröffentlicht: (2025)
Cut2Next: Generating Next Shot via In-Context Tuning
von: He, Jingwen, et al.
Veröffentlicht: (2025)
von: He, Jingwen, et al.
Veröffentlicht: (2025)
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
von: Pu, Yuandong, et al.
Veröffentlicht: (2025)
von: Pu, Yuandong, et al.
Veröffentlicht: (2025)
Lumina: Real-Time Mobile Neural Rendering by Exploiting Computational Redundancy
von: Feng, Yu, et al.
Veröffentlicht: (2025)
von: Feng, Yu, et al.
Veröffentlicht: (2025)
Point Transformer V3: Simpler, Faster, Stronger
von: Wu, Xiaoyang, et al.
Veröffentlicht: (2023)
von: Wu, Xiaoyang, et al.
Veröffentlicht: (2023)
The Lumina Project: CMB Optical Depth Fluctuations from Patchy Reionization
von: Smith, Aaron, et al.
Veröffentlicht: (2026)
von: Smith, Aaron, et al.
Veröffentlicht: (2026)
I-Max: Maximize the Resolution Potential of Pre-trained Rectified Flow Transformers with Projected Flow
von: Du, Ruoyi, et al.
Veröffentlicht: (2024)
von: Du, Ruoyi, et al.
Veröffentlicht: (2024)
Introducing the Lumina project: large-volume radiation-hydrodynamic simulations of the epochs of hydrogen and helium reionization
von: Zier, Oliver, et al.
Veröffentlicht: (2026)
von: Zier, Oliver, et al.
Veröffentlicht: (2026)
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction
von: Zhang, Xuying, et al.
Veröffentlicht: (2024)
von: Zhang, Xuying, et al.
Veröffentlicht: (2024)
GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
The Lumina Project: The Demographics of Active Galactic Nuclei from Quasars to Little Red Dots at $z\geq 3$
von: Shen, Xuejian, et al.
Veröffentlicht: (2026)
von: Shen, Xuejian, et al.
Veröffentlicht: (2026)
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
TaQ-DiT: Time-aware Quantization for Diffusion Transformers
von: Liu, Xinyan, et al.
Veröffentlicht: (2024)
von: Liu, Xinyan, et al.
Veröffentlicht: (2024)
LaVin-DiT: Large Vision Diffusion Transformer
von: Wang, Zhaoqing, et al.
Veröffentlicht: (2024)
von: Wang, Zhaoqing, et al.
Veröffentlicht: (2024)
Learning to Integrate Diffusion ODEs by Averaging the Derivatives
von: Liu, Wenze, et al.
Veröffentlicht: (2025)
von: Liu, Wenze, et al.
Veröffentlicht: (2025)
FD-DiT: Frequency Domain-Directed Diffusion Transformer for Low-Dose CT Reconstruction
von: Liu, Qiqing, et al.
Veröffentlicht: (2025)
von: Liu, Qiqing, et al.
Veröffentlicht: (2025)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
HQ-DiT: Efficient Diffusion Transformer with FP4 Hybrid Quantization
von: Liu, Wenxuan, et al.
Veröffentlicht: (2024)
von: Liu, Wenxuan, et al.
Veröffentlicht: (2024)
Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
von: Liu, Yuliang, et al.
Veröffentlicht: (2026)
von: Liu, Yuliang, et al.
Veröffentlicht: (2026)
Towards Efficient and Intelligent Laser Weeding: Method and Dataset for Weed Stem Detection
von: Liu, Dingning, et al.
Veröffentlicht: (2025)
von: Liu, Dingning, et al.
Veröffentlicht: (2025)
FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification
von: Yao, Jingfeng, et al.
Veröffentlicht: (2024)
von: Yao, Jingfeng, et al.
Veröffentlicht: (2024)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
von: Sun, Yasheng, et al.
Veröffentlicht: (2025)
von: Sun, Yasheng, et al.
Veröffentlicht: (2025)
Flow-GRPO: Training Flow Matching Models via Online RL
von: Liu, Jie, et al.
Veröffentlicht: (2025)
von: Liu, Jie, et al.
Veröffentlicht: (2025)
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
von: Ma, Teli, et al.
Veröffentlicht: (2026)
von: Ma, Teli, et al.
Veröffentlicht: (2026)
DiVE: DiT-based Video Generation with Enhanced Control
von: Jiang, Junpeng, et al.
Veröffentlicht: (2024)
von: Jiang, Junpeng, et al.
Veröffentlicht: (2024)
Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
von: Liu, Dingning, et al.
Veröffentlicht: (2024)
von: Liu, Dingning, et al.
Veröffentlicht: (2024)
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
Next Tokens Denoising for Speech Synthesis
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
GS-DiT: Advancing Video Generation with Pseudo 4D Gaussian Fields through Efficient Dense 3D Point Tracking
von: Bian, Weikang, et al.
Veröffentlicht: (2025)
von: Bian, Weikang, et al.
Veröffentlicht: (2025)
Transition Models: Rethinking the Generative Learning Objective
von: Wang, Zidong, et al.
Veröffentlicht: (2025)
von: Wang, Zidong, et al.
Veröffentlicht: (2025)
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
von: Li, Zhimin, et al.
Veröffentlicht: (2024)
von: Li, Zhimin, et al.
Veröffentlicht: (2024)
On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
von: Liu, Dongyang, et al.
Veröffentlicht: (2025) -
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
von: Qin, Qi, et al.
Veröffentlicht: (2025) -
Lumina
Veröffentlicht: (2017) -
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
von: Liu, Dongyang, et al.
Veröffentlicht: (2024) -
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
von: Gao, Peng, et al.
Veröffentlicht: (2024)