Vision Generalist Model: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ziyi, Rao, Yongming, Sun, Shuofeng, Liu, Xinrun, Wei, Yi, Yu, Xumin, Liu, Zuyan, Wang, Yanbo, Liu, Hongmin, Zhou, Jie, Lu, Jiwen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912425213689856
author Wang, Ziyi
Rao, Yongming
Sun, Shuofeng
Liu, Xinrun
Wei, Yi
Yu, Xumin
Liu, Zuyan
Wang, Yanbo
Liu, Hongmin
Zhou, Jie
Lu, Jiwen
author_facet Wang, Ziyi
Rao, Yongming
Sun, Shuofeng
Liu, Xinrun
Wei, Yi
Yu, Xumin
Liu, Zuyan
Wang, Yanbo
Liu, Hongmin
Zhou, Jie
Lu, Jiwen
contents Recently, we have witnessed the great success of the generalist model in natural language processing. The generalist model is a general framework trained with massive data and is able to process various downstream tasks simultaneously. Encouraged by their impressive performance, an increasing number of researchers are venturing into the realm of applying these models to computer vision tasks. However, the inputs and outputs of vision tasks are more diverse, and it is difficult to summarize them as a unified representation. In this paper, we provide a comprehensive overview of the vision generalist models, delving into their characteristics and capabilities within the field. First, we review the background, including the datasets, tasks, and benchmarks. Then, we dig into the design of frameworks that have been proposed in existing research, while also introducing the techniques employed to enhance their performance. To better help the researchers comprehend the area, we take a brief excursion into related domains, shedding light on their interconnections and potential synergies. To conclude, we provide some real-world application scenarios, undertake a thorough examination of the persistent challenges, and offer insights into possible directions for future research endeavors.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09954
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision Generalist Model: A Survey
Wang, Ziyi
Rao, Yongming
Sun, Shuofeng
Liu, Xinrun
Wei, Yi
Yu, Xumin
Liu, Zuyan
Wang, Yanbo
Liu, Hongmin
Zhou, Jie
Lu, Jiwen
Computer Vision and Pattern Recognition
Artificial Intelligence
Recently, we have witnessed the great success of the generalist model in natural language processing. The generalist model is a general framework trained with massive data and is able to process various downstream tasks simultaneously. Encouraged by their impressive performance, an increasing number of researchers are venturing into the realm of applying these models to computer vision tasks. However, the inputs and outputs of vision tasks are more diverse, and it is difficult to summarize them as a unified representation. In this paper, we provide a comprehensive overview of the vision generalist models, delving into their characteristics and capabilities within the field. First, we review the background, including the datasets, tasks, and benchmarks. Then, we dig into the design of frameworks that have been proposed in existing research, while also introducing the techniques employed to enhance their performance. To better help the researchers comprehend the area, we take a brief excursion into related domains, shedding light on their interconnections and potential synergies. To conclude, we provide some real-world application scenarios, undertake a thorough examination of the persistent challenges, and offer insights into possible directions for future research endeavors.
title Vision Generalist Model: A Survey
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.09954