Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Shufan, Gu, Jiuxiang, Liu, Kangning, Lin, Zhe, Wei, Zijun, Grover, Aditya, Kuen, Jason
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909803832410112
author Li, Shufan
Gu, Jiuxiang
Liu, Kangning
Lin, Zhe
Wei, Zijun
Grover, Aditya
Kuen, Jason
author_facet Li, Shufan
Gu, Jiuxiang
Liu, Kangning
Lin, Zhe
Wei, Zijun
Grover, Aditya
Kuen, Jason
contents We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks and low-resolution image generation, Lavida-O presents a single framework that enables image-level understanding, object grounding, image editing, and high-resolution (1024px) text-to-image synthesis. Lavida-O incorporates a novel Elastic Mixture-of-Transformers (Elastic-MoT) architecture that couples a lightweight generation branch with a larger understanding branch, supported by token compression, universal text conditioning and stratified sampling for efficient and high-quality generation. Lavida-O further incorporates planning and iterative self-reflection in image generation and editing tasks, seamlessly boosting generation quality with its understanding capabilities. Lavida-O achieves state-of-the-art performance on a wide range of benchmarks including RefCOCO object grounding, GenEval text-to-image generation, and ImgEdit image editing, outperforming existing autoregressive models and continuous diffusion models such as Qwen2.5-VL and FluxKontext-dev, while offering considerable speedup at inference. These advances establish Lavida-O as a new paradigm for scalable multimodal reasoning and generation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19244
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
Li, Shufan
Gu, Jiuxiang
Liu, Kangning
Lin, Zhe
Wei, Zijun
Grover, Aditya
Kuen, Jason
Computer Vision and Pattern Recognition
We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks and low-resolution image generation, Lavida-O presents a single framework that enables image-level understanding, object grounding, image editing, and high-resolution (1024px) text-to-image synthesis. Lavida-O incorporates a novel Elastic Mixture-of-Transformers (Elastic-MoT) architecture that couples a lightweight generation branch with a larger understanding branch, supported by token compression, universal text conditioning and stratified sampling for efficient and high-quality generation. Lavida-O further incorporates planning and iterative self-reflection in image generation and editing tasks, seamlessly boosting generation quality with its understanding capabilities. Lavida-O achieves state-of-the-art performance on a wide range of benchmarks including RefCOCO object grounding, GenEval text-to-image generation, and ImgEdit image editing, outperforming existing autoregressive models and continuous diffusion models such as Qwen2.5-VL and FluxKontext-dev, while offering considerable speedup at inference. These advances establish Lavida-O as a new paradigm for scalable multimodal reasoning and generation.
title Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.19244