VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915182101397504 |
|---|---|
| author | Wu, Yecheng Zhang, Zhuoyang Chen, Junyu Tang, Haotian Li, Dacheng Fang, Yunhao Zhu, Ligeng Xie, Enze Yin, Hongxu Yi, Li Han, Song Lu, Yao |
| author_facet | Wu, Yecheng Zhang, Zhuoyang Chen, Junyu Tang, Haotian Li, Dacheng Fang, Yunhao Zhu, Ligeng Xie, Enze Yin, Hongxu Yi, Li Han, Song Lu, Yao |
| contents | VILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for understanding and generating visual content, which can lead to misalignment and increased complexity. In contrast, VILA-U employs a single autoregressive next-token prediction framework for both tasks, eliminating the need for additional components like diffusion models. This approach not only simplifies the model but also achieves near state-of-the-art performance in visual language understanding and generation. The success of VILA-U is attributed to two main factors: the unified vision tower that aligns discrete visual tokens with textual inputs during pretraining, which enhances visual perception, and autoregressive image generation can achieve similar quality as diffusion models with high-quality dataset. This allows VILA-U to perform comparably to more complex models using a fully token-based autoregressive framework. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_04429 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation Wu, Yecheng Zhang, Zhuoyang Chen, Junyu Tang, Haotian Li, Dacheng Fang, Yunhao Zhu, Ligeng Xie, Enze Yin, Hongxu Yi, Li Han, Song Lu, Yao Computer Vision and Pattern Recognition Machine Learning VILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for understanding and generating visual content, which can lead to misalignment and increased complexity. In contrast, VILA-U employs a single autoregressive next-token prediction framework for both tasks, eliminating the need for additional components like diffusion models. This approach not only simplifies the model but also achieves near state-of-the-art performance in visual language understanding and generation. The success of VILA-U is attributed to two main factors: the unified vision tower that aligns discrete visual tokens with textual inputs during pretraining, which enhances visual perception, and autoregressive image generation can achieve similar quality as diffusion models with high-quality dataset. This allows VILA-U to perform comparably to more complex models using a fully token-based autoregressive framework. |
| title | VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2409.04429 |