VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wu, Yecheng, Zhang, Zhuoyang, Chen, Junyu, Tang, Haotian, Li, Dacheng, Fang, Yunhao, Zhu, Ligeng, Xie, Enze, Yin, Hongxu, Yi, Li, Han, Song, Lu, Yao
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915182101397504
author Wu, Yecheng
Zhang, Zhuoyang
Chen, Junyu
Tang, Haotian
Li, Dacheng
Fang, Yunhao
Zhu, Ligeng
Xie, Enze
Yin, Hongxu
Yi, Li
Han, Song
Lu, Yao
author_facet Wu, Yecheng
Zhang, Zhuoyang
Chen, Junyu
Tang, Haotian
Li, Dacheng
Fang, Yunhao
Zhu, Ligeng
Xie, Enze
Yin, Hongxu
Yi, Li
Han, Song
Lu, Yao
contents VILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for understanding and generating visual content, which can lead to misalignment and increased complexity. In contrast, VILA-U employs a single autoregressive next-token prediction framework for both tasks, eliminating the need for additional components like diffusion models. This approach not only simplifies the model but also achieves near state-of-the-art performance in visual language understanding and generation. The success of VILA-U is attributed to two main factors: the unified vision tower that aligns discrete visual tokens with textual inputs during pretraining, which enhances visual perception, and autoregressive image generation can achieve similar quality as diffusion models with high-quality dataset. This allows VILA-U to perform comparably to more complex models using a fully token-based autoregressive framework.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04429
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Wu, Yecheng
Zhang, Zhuoyang
Chen, Junyu
Tang, Haotian
Li, Dacheng
Fang, Yunhao
Zhu, Ligeng
Xie, Enze
Yin, Hongxu
Yi, Li
Han, Song
Lu, Yao
Computer Vision and Pattern Recognition
Machine Learning
VILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for understanding and generating visual content, which can lead to misalignment and increased complexity. In contrast, VILA-U employs a single autoregressive next-token prediction framework for both tasks, eliminating the need for additional components like diffusion models. This approach not only simplifies the model but also achieves near state-of-the-art performance in visual language understanding and generation. The success of VILA-U is attributed to two main factors: the unified vision tower that aligns discrete visual tokens with textual inputs during pretraining, which enhances visual perception, and autoregressive image generation can achieve similar quality as diffusion models with high-quality dataset. This allows VILA-U to perform comparably to more complex models using a fully token-based autoregressive framework.
title VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2409.04429