Guardado en:
Detalles Bibliográficos
Autores principales: Pan, Kaihang, Lin, Wang, Yue, Zhongqi, Ao, Tenglong, Jia, Liyu, Zhao, Wei, Li, Juncheng, Tang, Siliang, Zhang, Hanwang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2504.14666
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912337796005888
author Pan, Kaihang
Lin, Wang
Yue, Zhongqi
Ao, Tenglong
Jia, Liyu
Zhao, Wei
Li, Juncheng
Tang, Siliang
Zhang, Hanwang
author_facet Pan, Kaihang
Lin, Wang
Yue, Zhongqi
Ao, Tenglong
Jia, Liyu
Zhao, Wei
Li, Juncheng
Tang, Siliang
Zhang, Hanwang
contents Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation by combining LLM and diffusion models, the state-of-the-art in each task, respectively. Existing approaches rely on spatial visual tokens, where image patches are encoded and arranged according to a spatial order (e.g., raster scan). However, we show that spatial tokens lack the recursive structure inherent to languages, hence form an impossible language for LLM to master. In this paper, we build a proper visual language by leveraging diffusion timesteps to learn discrete, recursive visual tokens. Our proposed tokens recursively compensate for the progressive attribute loss in noisy images as timesteps increase, enabling the diffusion model to reconstruct the original image at any timestep. This approach allows us to effectively integrate the strengths of LLMs in autoregressive reasoning and diffusion models in precise image generation, achieving seamless multimodal comprehension and generation within a unified framework. Extensive experiments show that we achieve superior performance for multimodal comprehension and generation simultaneously compared with other MLLMs. Project Page: https://DDT-LLaMA.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14666
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
Pan, Kaihang
Lin, Wang
Yue, Zhongqi
Ao, Tenglong
Jia, Liyu
Zhao, Wei
Li, Juncheng
Tang, Siliang
Zhang, Hanwang
Computer Vision and Pattern Recognition
Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation by combining LLM and diffusion models, the state-of-the-art in each task, respectively. Existing approaches rely on spatial visual tokens, where image patches are encoded and arranged according to a spatial order (e.g., raster scan). However, we show that spatial tokens lack the recursive structure inherent to languages, hence form an impossible language for LLM to master. In this paper, we build a proper visual language by leveraging diffusion timesteps to learn discrete, recursive visual tokens. Our proposed tokens recursively compensate for the progressive attribute loss in noisy images as timesteps increase, enabling the diffusion model to reconstruct the original image at any timestep. This approach allows us to effectively integrate the strengths of LLMs in autoregressive reasoning and diffusion models in precise image generation, achieving seamless multimodal comprehension and generation within a unified framework. Extensive experiments show that we achieve superior performance for multimodal comprehension and generation simultaneously compared with other MLLMs. Project Page: https://DDT-LLaMA.github.io/.
title Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.14666