Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Xin, Yi, Qin, Qi, Luo, Siqi, Zhu, Kaiwen, Yan, Juncheng, Tai, Yan, Lei, Jiayi, Cao, Yuewen, Wang, Keqi, Wang, Yibin, Bai, Jinbin, Yu, Qian, Jiang, Dengyang, Pu, Yuandong, Chen, Haoxing, Zhuo, Le, He, Junjun, Luo, Gen, Li, Tianbin, Hu, Ming, Ye, Jin, Ye, Shenglong, Zhang, Bo, Xu, Chang, Wang, Wenhai, Li, Hongsheng, Zhai, Guangtao, Xue, Tianfan, Fu, Bin, Liu, Xiaohong, Qiao, Yu, Liu, Yihao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
by: Pu, Yuandong, et al.
Published: (2025)
by: Pu, Yuandong, et al.
Published: (2025)
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
by: Xin, Yi, et al.
Published: (2025)
by: Xin, Yi, et al.
Published: (2025)
OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM
by: Hu, Yutao, et al.
Published: (2024)
by: Hu, Yutao, et al.
Published: (2024)
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
“Store Strategy”: A New Omni‐Channel Strategy in Community Group Buying
by: Nana Zhang, et al.
Published: (2024)
by: Nana Zhang, et al.
Published: (2024)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
by: Liu, Dongyang, et al.
Published: (2025)
by: Liu, Dongyang, et al.
Published: (2025)
Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models
by: Luo, Siqi, et al.
Published: (2026)
by: Luo, Siqi, et al.
Published: (2026)
Lumina
Published: (2017)
Published: (2017)
Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings
by: Wei, Xingguang, et al.
Published: (2025)
by: Wei, Xingguang, et al.
Published: (2025)
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
by: Qin, Qi, et al.
Published: (2025)
by: Qin, Qi, et al.
Published: (2025)
Counting Permutations in $S_{2n}$ and $S_{2n+1}$
by: Luo, Yuewen
Published: (2024)
by: Luo, Yuewen
Published: (2024)
Elastic constant ratio for fatigue evaluation on rubber isolators
by: Robert Keqi Luo
Published: (2024)
by: Robert Keqi Luo
Published: (2024)
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
by: Xin, Yi, et al.
Published: (2025)
by: Xin, Yi, et al.
Published: (2025)
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
by: Gao, Peng, et al.
Published: (2024)
by: Gao, Peng, et al.
Published: (2024)
Preconditioned Inexact Stochastic ADMM for Deep Model
by: Zhou, Shenglong, et al.
Published: (2025)
by: Zhou, Shenglong, et al.
Published: (2025)
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
by: Liu, Dongyang, et al.
Published: (2024)
by: Liu, Dongyang, et al.
Published: (2024)
ICEPOP-2018 project MOO PARSIVEL
by: Kyungpook National University
Published: (2025)
by: Kyungpook National University
Published: (2025)
SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding
by: Chen, Ying, et al.
Published: (2024)
by: Chen, Ying, et al.
Published: (2024)
HSD: Training-Free Acceleration for Document Parsing Vision-Language Model with Hierarchical Speculative Decoding
by: Liao, Wenhui, et al.
Published: (2026)
by: Liao, Wenhui, et al.
Published: (2026)
BADM: Batch ADMM for Deep Learning
by: Wang, Ouya, et al.
Published: (2024)
by: Wang, Ouya, et al.
Published: (2024)
SAM-Med3D-MoE: Towards a Non-Forgetting Segment Anything Model via Mixture of Experts for 3D Medical Image Segmentation
by: Wang, Guoan, et al.
Published: (2024)
by: Wang, Guoan, et al.
Published: (2024)
Docopilot: Improving Multimodal Models for Document-Level Understanding
by: Duan, Yuchen, et al.
Published: (2025)
by: Duan, Yuchen, et al.
Published: (2025)
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
Generic linear convergence for algorithms of non-linear least squares over smooth varieties
by: Hu, Shenglong, et al.
Published: (2025)
by: Hu, Shenglong, et al.
Published: (2025)
ArchCAD-400K: A Large-Scale CAD drawings Dataset and New Baseline for Panoptic Symbol Spotting
by: Luo, Ruifeng, et al.
Published: (2025)
by: Luo, Ruifeng, et al.
Published: (2025)
Efficacy and Safety of Biologics for Chronic Rhinosinusitis With Nasal Polyps: A Meta‐Analysis of Real‐World Evidence
by: Shiru Cai, et al.
Published: (2025)
by: Shiru Cai, et al.
Published: (2025)
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
by: Yan, Yibin, et al.
Published: (2026)
by: Yan, Yibin, et al.
Published: (2026)
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
by: Xie, Tianyu, et al.
Published: (2026)
by: Xie, Tianyu, et al.
Published: (2026)
Uncertainty-Adjusted Sorting for Asset Pricing with Machine Learning
by: Liu, Yan, et al.
Published: (2026)
by: Liu, Yan, et al.
Published: (2026)
Chiral spin state and nematic ferromagnet in the spin-1 Kitaev-$Γ$ model
by: Luo, Qiang, et al.
Published: (2024)
by: Luo, Qiang, et al.
Published: (2024)
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
by: Li, Qingyun, et al.
Published: (2024)
by: Li, Qingyun, et al.
Published: (2024)
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
by: Gao, Zhangwei, et al.
Published: (2024)
by: Gao, Zhangwei, et al.
Published: (2024)
Managing the Catalyst Energy Landscape: Continuous Fermi‐Level Tuning Eliminates the Overpotential Penalty in Organic Electrooxidation
by: Liping Wang, et al.
Published: (2026)
by: Liping Wang, et al.
Published: (2026)
FedGiA: An Efficient Hybrid Algorithm for Federated Learning
by: Zhou, Shenglong, et al.
Published: (2022)
by: Zhou, Shenglong, et al.
Published: (2022)
Lumina: Real-Time Mobile Neural Rendering by Exploiting Computational Redundancy
by: Feng, Yu, et al.
Published: (2025)
by: Feng, Yu, et al.
Published: (2025)
EdMOO: One Approach to a Multimedia Collaborative Environment.
by: Holkner, Bernard
Published: (1996)
by: Holkner, Bernard
Published: (1996)
An adaptive symplectic integrator for gravitational dynamics
by: Ye, Keqi, et al.
Published: (2025)
by: Ye, Keqi, et al.
Published: (2025)
H4K12 Lactylation Activated‐ Spp1 in Reprogrammed Microglia Improves Functional Recovery After Spinal Cord Injury
by: Xiaokun Wang, et al.
Published: (2025)
by: Xiaokun Wang, et al.
Published: (2025)
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
by: Ren, Yiming, et al.
Published: (2025)
by: Ren, Yiming, et al.
Published: (2025)
Interpretable Interaction Modeling for Trajectory Prediction via Agent Selection and Physical Coefficient
by: Huang, Shiji, et al.
Published: (2024)
by: Huang, Shiji, et al.
Published: (2024)
Similar Items
-
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
by: Pu, Yuandong, et al.
Published: (2025) -
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
by: Xin, Yi, et al.
Published: (2025) -
OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM
by: Hu, Yutao, et al.
Published: (2024) -
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
by: Zhuo, Le, et al.
Published: (2024) -
“Store Strategy”: A New Omni‐Channel Strategy in Community Group Buying
by: Nana Zhang, et al.
Published: (2024)