LaVin-DiT: Large Vision Diffusion Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhaoqing, Xia, Xiaobo, Chen, Runnan, Yu, Dongdong, Wang, Changhu, Gong, Mingming, Liu, Tongliang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909526527049728
author Wang, Zhaoqing
Xia, Xiaobo
Chen, Runnan
Yu, Dongdong
Wang, Changhu
Gong, Mingming
Liu, Tongliang
author_facet Wang, Zhaoqing
Xia, Xiaobo
Chen, Runnan
Yu, Dongdong
Wang, Changhu
Gong, Mingming
Liu, Tongliang
contents This paper presents the Large Vision Diffusion Transformer (LaVin-DiT), a scalable and unified foundation model designed to tackle over 20 computer vision tasks in a generative framework. Unlike existing large vision models directly adapted from natural language processing architectures, which rely on less efficient autoregressive techniques and disrupt spatial relationships essential for vision data, LaVin-DiT introduces key innovations to optimize generative performance for vision tasks. First, to address the high dimensionality of visual data, we incorporate a spatial-temporal variational autoencoder that encodes data into a continuous latent space. Second, for generative modeling, we develop a joint diffusion transformer that progressively produces vision outputs. Third, for unified multi-task training, in-context learning is implemented. Input-target pairs serve as task context, which guides the diffusion transformer to align outputs with specific tasks within the latent space. During inference, a task-specific context set and test data as queries allow LaVin-DiT to generalize across tasks without fine-tuning. Trained on extensive vision datasets, the model is scaled from 0.1B to 3.4B parameters, demonstrating substantial scalability and state-of-the-art performance across diverse vision tasks. This work introduces a novel pathway for large vision foundation models, underscoring the promising potential of diffusion transformers. The code and models are available.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11505
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LaVin-DiT: Large Vision Diffusion Transformer
Wang, Zhaoqing
Xia, Xiaobo
Chen, Runnan
Yu, Dongdong
Wang, Changhu
Gong, Mingming
Liu, Tongliang
Computer Vision and Pattern Recognition
This paper presents the Large Vision Diffusion Transformer (LaVin-DiT), a scalable and unified foundation model designed to tackle over 20 computer vision tasks in a generative framework. Unlike existing large vision models directly adapted from natural language processing architectures, which rely on less efficient autoregressive techniques and disrupt spatial relationships essential for vision data, LaVin-DiT introduces key innovations to optimize generative performance for vision tasks. First, to address the high dimensionality of visual data, we incorporate a spatial-temporal variational autoencoder that encodes data into a continuous latent space. Second, for generative modeling, we develop a joint diffusion transformer that progressively produces vision outputs. Third, for unified multi-task training, in-context learning is implemented. Input-target pairs serve as task context, which guides the diffusion transformer to align outputs with specific tasks within the latent space. During inference, a task-specific context set and test data as queries allow LaVin-DiT to generalize across tasks without fine-tuning. Trained on extensive vision datasets, the model is scaled from 0.1B to 3.4B parameters, demonstrating substantial scalability and state-of-the-art performance across diverse vision tasks. This work introduces a novel pathway for large vision foundation models, underscoring the promising potential of diffusion transformers. The code and models are available.
title LaVin-DiT: Large Vision Diffusion Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.11505