Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Shenghao, Zu, Wenqiang, Zhao, Mingyang, Su, Duo, Liu, Shilong, Shi, Ruohua, Li, Guoqi, Zhang, Shanghang, Ma, Lei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913567825985536
author Xie, Shenghao
Zu, Wenqiang
Zhao, Mingyang
Su, Duo
Liu, Shilong
Shi, Ruohua
Li, Guoqi
Zhang, Shanghang
Ma, Lei
author_facet Xie, Shenghao
Zu, Wenqiang
Zhao, Mingyang
Su, Duo
Liu, Shilong
Shi, Ruohua
Li, Guoqi
Zhang, Shanghang
Ma, Lei
contents Autoregression in large language models (LLMs) has shown impressive scalability by unifying all language tasks into the next token prediction paradigm. Recently, there is a growing interest in extending this success to vision foundation models. In this survey, we review the recent advances and discuss future directions for autoregressive vision foundation models. First, we present the trend for next generation of vision foundation models, i.e., unifying both understanding and generation in vision tasks. We then analyze the limitations of existing vision foundation models, and present a formal definition of autoregression with its advantages. Later, we categorize autoregressive vision foundation models from their vision tokenizers and autoregression backbones. Finally, we discuss several promising research challenges and directions. To the best of our knowledge, this is the first survey to comprehensively summarize autoregressive vision foundation models under the trend of unifying understanding and generation. A collection of related resources is available at https://github.com/EmmaSRH/ARVFM.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22217
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective
Xie, Shenghao
Zu, Wenqiang
Zhao, Mingyang
Su, Duo
Liu, Shilong
Shi, Ruohua
Li, Guoqi
Zhang, Shanghang
Ma, Lei
Computer Vision and Pattern Recognition
Autoregression in large language models (LLMs) has shown impressive scalability by unifying all language tasks into the next token prediction paradigm. Recently, there is a growing interest in extending this success to vision foundation models. In this survey, we review the recent advances and discuss future directions for autoregressive vision foundation models. First, we present the trend for next generation of vision foundation models, i.e., unifying both understanding and generation in vision tasks. We then analyze the limitations of existing vision foundation models, and present a formal definition of autoregression with its advantages. Later, we categorize autoregressive vision foundation models from their vision tokenizers and autoregression backbones. Finally, we discuss several promising research challenges and directions. To the best of our knowledge, this is the first survey to comprehensively summarize autoregressive vision foundation models under the trend of unifying understanding and generation. A collection of related resources is available at https://github.com/EmmaSRH/ARVFM.
title Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.22217