Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Peng, Zhuo, Le, Liu, Dongyang, Du, Ruoyi, Luo, Xu, Qiu, Longtian, Zhang, Yuhang, Lin, Chen, Huang, Rongjie, Geng, Shijie, Zhang, Renrui, Xi, Junlin, Shao, Wenqi, Jiang, Zhengkai, Yang, Tianshuo, Ye, Weicai, Tong, He, He, Jingwen, Qiao, Yu, Li, Hongsheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
di: Zhuo, Le, et al.
Pubblicazione: (2024)
di: Zhuo, Le, et al.
Pubblicazione: (2024)
I-Max: Maximize the Resolution Potential of Pre-trained Rectified Flow Transformers with Projected Flow
di: Du, Ruoyi, et al.
Pubblicazione: (2024)
di: Du, Ruoyi, et al.
Pubblicazione: (2024)
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
di: Qin, Qi, et al.
Pubblicazione: (2025)
di: Qin, Qi, et al.
Pubblicazione: (2025)
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
di: Liu, Dongyang, et al.
Pubblicazione: (2024)
di: Liu, Dongyang, et al.
Pubblicazione: (2024)
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
di: Liu, Dongyang, et al.
Pubblicazione: (2024)
di: Liu, Dongyang, et al.
Pubblicazione: (2024)
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
di: Xin, Yi, et al.
Pubblicazione: (2025)
di: Xin, Yi, et al.
Pubblicazione: (2025)
SGTR+: End-to-end Scene Graph Generation with Transformer
di: Li, Rongjie, et al.
Pubblicazione: (2024)
di: Li, Rongjie, et al.
Pubblicazione: (2024)
ViTAR: Vision Transformer with Any Resolution
di: Fan, Qihang, et al.
Pubblicazione: (2024)
di: Fan, Qihang, et al.
Pubblicazione: (2024)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
di: Gao, Peng, et al.
Pubblicazione: (2021)
di: Gao, Peng, et al.
Pubblicazione: (2021)
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
di: Xin, Yi, et al.
Pubblicazione: (2025)
di: Xin, Yi, et al.
Pubblicazione: (2025)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
di: Liu, Dongyang, et al.
Pubblicazione: (2025)
di: Liu, Dongyang, et al.
Pubblicazione: (2025)
MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any Resolution
di: Liu, Wenzhuo, et al.
Pubblicazione: (2024)
di: Liu, Wenzhuo, et al.
Pubblicazione: (2024)
Lumina
Pubblicazione: (2017)
Pubblicazione: (2017)
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
di: Pu, Yuandong, et al.
Pubblicazione: (2025)
di: Pu, Yuandong, et al.
Pubblicazione: (2025)
Eliminating VAE for Fast and High-Resolution Generative Detail Restoration
di: Wang, Yan, et al.
Pubblicazione: (2026)
di: Wang, Yan, et al.
Pubblicazione: (2026)
MonoDETR: Depth-guided Transformer for Monocular 3D Object Detection
di: Zhang, Renrui, et al.
Pubblicazione: (2022)
di: Zhang, Renrui, et al.
Pubblicazione: (2022)
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
di: Huang, Victor Shea-Jay, et al.
Pubblicazione: (2025)
di: Huang, Victor Shea-Jay, et al.
Pubblicazione: (2025)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
di: Lin, Weifeng, et al.
Pubblicazione: (2024)
di: Lin, Weifeng, et al.
Pubblicazione: (2024)
Lumina: Real-Time Mobile Neural Rendering by Exploiting Computational Redundancy
di: Feng, Yu, et al.
Pubblicazione: (2025)
di: Feng, Yu, et al.
Pubblicazione: (2025)
TerDiT: Ternary Diffusion Models with Transformers
di: Lu, Xudong, et al.
Pubblicazione: (2024)
di: Lu, Xudong, et al.
Pubblicazione: (2024)
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
di: Qiu, Longtian, et al.
Pubblicazione: (2024)
di: Qiu, Longtian, et al.
Pubblicazione: (2024)
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
di: Ning, Shan, et al.
Pubblicazione: (2026)
di: Ning, Shan, et al.
Pubblicazione: (2026)
Prediction of Bridge Structural Response Based on Nonstationary Transformer
di: Qing Li, et al.
Pubblicazione: (2025)
di: Qing Li, et al.
Pubblicazione: (2025)
Optimal Transcoding Resolution Prediction for Efficient Per-Title Bitrate Ladder Estimation
di: Yang, Jinhai, et al.
Pubblicazione: (2024)
di: Yang, Jinhai, et al.
Pubblicazione: (2024)
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
di: Cheng, Yifan, et al.
Pubblicazione: (2025)
di: Cheng, Yifan, et al.
Pubblicazione: (2025)
TAPTR: Tracking Any Point with Transformers as Detection
di: Li, Hongyang, et al.
Pubblicazione: (2024)
di: Li, Hongyang, et al.
Pubblicazione: (2024)
Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer
di: Yang, Zhuoyi, et al.
Pubblicazione: (2024)
di: Yang, Zhuoyi, et al.
Pubblicazione: (2024)
Improved Adversarial Diffusion Compression for Real-World Video Super-Resolution
di: Chen, Bin, et al.
Pubblicazione: (2026)
di: Chen, Bin, et al.
Pubblicazione: (2026)
Anomaly Detection in Industrial Control Systems Based on Cross-Domain Representation Learning
di: Zhan, Dongyang, et al.
Pubblicazione: (2025)
di: Zhan, Dongyang, et al.
Pubblicazione: (2025)
Beyond the Next Port: A Multi-Task Transformer for Forecasting Future Voyage Segment Durations
di: Liu, Nairui, et al.
Pubblicazione: (2026)
di: Liu, Nairui, et al.
Pubblicazione: (2026)
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
di: Li, Yiheng, et al.
Pubblicazione: (2026)
di: Li, Yiheng, et al.
Pubblicazione: (2026)
Tracking and Segmenting Anything in Any Modality
di: Zhang, Tianlu, et al.
Pubblicazione: (2025)
di: Zhang, Tianlu, et al.
Pubblicazione: (2025)
Audio-Visual Cross-Modal Compression for Generative Face Video Coding
di: Xu, Youmin, et al.
Pubblicazione: (2025)
di: Xu, Youmin, et al.
Pubblicazione: (2025)
High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer
di: Zheng, Shen, et al.
Pubblicazione: (2025)
di: Zheng, Shen, et al.
Pubblicazione: (2025)
Simba: Towards High-Fidelity and Geometrically-Consistent Point Cloud Completion via Transformation Diffusion
di: Zhang, Lirui, et al.
Pubblicazione: (2025)
di: Zhang, Lirui, et al.
Pubblicazione: (2025)
DA-DPO: Cost-efficient Difficulty-aware Preference Optimization for Reducing MLLM Hallucinations
di: Qiu, Longtian, et al.
Pubblicazione: (2026)
di: Qiu, Longtian, et al.
Pubblicazione: (2026)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
di: He, Xuanhua, et al.
Pubblicazione: (2025)
di: He, Xuanhua, et al.
Pubblicazione: (2025)
Judge Anything: MLLM as a Judge Across Any Modality
di: Pu, Shu, et al.
Pubblicazione: (2025)
di: Pu, Shu, et al.
Pubblicazione: (2025)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
di: Chen, Junyi, et al.
Pubblicazione: (2026)
di: Chen, Junyi, et al.
Pubblicazione: (2026)
From Sketch to Fresco: Efficient Diffusion Transformer with Progressive Resolution
di: Zheng, Shikang, et al.
Pubblicazione: (2026)
di: Zheng, Shikang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
di: Zhuo, Le, et al.
Pubblicazione: (2024) -
I-Max: Maximize the Resolution Potential of Pre-trained Rectified Flow Transformers with Projected Flow
di: Du, Ruoyi, et al.
Pubblicazione: (2024) -
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
di: Qin, Qi, et al.
Pubblicazione: (2025) -
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
di: Liu, Dongyang, et al.
Pubblicazione: (2024) -
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
di: Liu, Dongyang, et al.
Pubblicazione: (2024)