Delving into Multi-modal Multi-task Foundation Models for Road Scene Understanding: From Learning Paradigm Perspectives

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Sheng, Chen, Wei, Tian, Wanxin, Liu, Rui, Hou, Luanxuan, Zhang, Xiubao, Shen, Haifeng, Wu, Ruiqi, Geng, Shuyi, Zhou, Yi, Shao, Ling, Yang, Yi, Gao, Bojun, Li, Qun, Wu, Guobin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909210743144448
author Luo, Sheng
Chen, Wei
Tian, Wanxin
Liu, Rui
Hou, Luanxuan
Zhang, Xiubao
Shen, Haifeng
Wu, Ruiqi
Geng, Shuyi
Zhou, Yi
Shao, Ling
Yang, Yi
Gao, Bojun
Li, Qun
Wu, Guobin
author_facet Luo, Sheng
Chen, Wei
Tian, Wanxin
Liu, Rui
Hou, Luanxuan
Zhang, Xiubao
Shen, Haifeng
Wu, Ruiqi
Geng, Shuyi
Zhou, Yi
Shao, Ling
Yang, Yi
Gao, Bojun
Li, Qun
Wu, Guobin
contents Foundation models have indeed made a profound impact on various fields, emerging as pivotal components that significantly shape the capabilities of intelligent systems. In the context of intelligent vehicles, leveraging the power of foundation models has proven to be transformative, offering notable advancements in visual understanding. Equipped with multi-modal and multi-task learning capabilities, multi-modal multi-task visual understanding foundation models (MM-VUFMs) effectively process and fuse data from diverse modalities and simultaneously handle various driving-related tasks with powerful adaptability, contributing to a more holistic understanding of the surrounding scene. In this survey, we present a systematic analysis of MM-VUFMs specifically designed for road scenes. Our objective is not only to provide a comprehensive overview of common practices, referring to task-specific models, unified multi-modal models, unified multi-task models, and foundation model prompting techniques, but also to highlight their advanced capabilities in diverse learning paradigms. These paradigms include open-world understanding, efficient transfer for road scenes, continual learning, interactive and generative capability. Moreover, we provide insights into key challenges and future trends, such as closed-loop driving systems, interpretability, embodied driving agents, and world models. To facilitate researchers in staying abreast of the latest developments in MM-VUFMs for road scenes, we have established a continuously updated repository at https://github.com/rolsheng/MM-VUFM4DS
format Preprint
id arxiv_https___arxiv_org_abs_2402_02968
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Delving into Multi-modal Multi-task Foundation Models for Road Scene Understanding: From Learning Paradigm Perspectives
Luo, Sheng
Chen, Wei
Tian, Wanxin
Liu, Rui
Hou, Luanxuan
Zhang, Xiubao
Shen, Haifeng
Wu, Ruiqi
Geng, Shuyi
Zhou, Yi
Shao, Ling
Yang, Yi
Gao, Bojun
Li, Qun
Wu, Guobin
Computer Vision and Pattern Recognition
Machine Learning
Foundation models have indeed made a profound impact on various fields, emerging as pivotal components that significantly shape the capabilities of intelligent systems. In the context of intelligent vehicles, leveraging the power of foundation models has proven to be transformative, offering notable advancements in visual understanding. Equipped with multi-modal and multi-task learning capabilities, multi-modal multi-task visual understanding foundation models (MM-VUFMs) effectively process and fuse data from diverse modalities and simultaneously handle various driving-related tasks with powerful adaptability, contributing to a more holistic understanding of the surrounding scene. In this survey, we present a systematic analysis of MM-VUFMs specifically designed for road scenes. Our objective is not only to provide a comprehensive overview of common practices, referring to task-specific models, unified multi-modal models, unified multi-task models, and foundation model prompting techniques, but also to highlight their advanced capabilities in diverse learning paradigms. These paradigms include open-world understanding, efficient transfer for road scenes, continual learning, interactive and generative capability. Moreover, we provide insights into key challenges and future trends, such as closed-loop driving systems, interpretability, embodied driving agents, and world models. To facilitate researchers in staying abreast of the latest developments in MM-VUFMs for road scenes, we have established a continuously updated repository at https://github.com/rolsheng/MM-VUFM4DS
title Delving into Multi-modal Multi-task Foundation Models for Road Scene Understanding: From Learning Paradigm Perspectives
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2402.02968