Delving into Multi-modal Multi-task Foundation Models for Road Scene Understanding: From Learning Paradigm Perspectives
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Sheng, Chen, Wei, Tian, Wanxin, Liu, Rui, Hou, Luanxuan, Zhang, Xiubao, Shen, Haifeng, Wu, Ruiqi, Geng, Shuyi, Zhou, Yi, Shao, Ling, Yang, Yi, Gao, Bojun, Li, Qun, Wu, Guobin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey
by: Fu, Ao, et al.
Published: (2024)
by: Fu, Ao, et al.
Published: (2024)
Multi‐modality multiorgan image segmentation using continual learning with enhanced hard attention to the task
by: Ming‐Long Wu, et al.
Published: (2025)
by: Ming‐Long Wu, et al.
Published: (2025)
Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking
by: Zhang, Haonan, et al.
Published: (2025)
by: Zhang, Haonan, et al.
Published: (2025)
360+x: A Panoptic Multi-modal Scene Understanding Dataset
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
SocialGesture: Delving into Multi-person Gesture Understanding
by: Cao, Xu, et al.
Published: (2025)
by: Cao, Xu, et al.
Published: (2025)
Multi-modal Multi-task Pre-training for Improved Point Cloud Understanding
by: Liu, Liwen, et al.
Published: (2025)
by: Liu, Liwen, et al.
Published: (2025)
$M^3EL$: A Multi-task Multi-topic Dataset for Multi-modal Entity Linking
by: Wang, Fang, et al.
Published: (2024)
by: Wang, Fang, et al.
Published: (2024)
Delve into the Applicability of Advanced Optimizers for Multi-Task Learning
by: Zhou, Zhipeng, et al.
Published: (2026)
by: Zhou, Zhipeng, et al.
Published: (2026)
MMSense: Adapting Vision-based Foundation Model for Multi-task Multi-modal Wireless Sensing
by: Li, Zhizhen, et al.
Published: (2025)
by: Li, Zhizhen, et al.
Published: (2025)
Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
by: Wu, Zichen, et al.
Published: (2024)
by: Wu, Zichen, et al.
Published: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
by: Zhang, Ming, et al.
Published: (2024)
by: Zhang, Ming, et al.
Published: (2024)
Multi-modal Deepfake Detection and Localization with FPN-Transformer
by: Zheng, Chende, et al.
Published: (2025)
by: Zheng, Chende, et al.
Published: (2025)
Efficient RGB-D Scene Understanding via Multi-task Adaptive Learning and Cross-dimensional Feature Guidance
by: Sun, Guodong, et al.
Published: (2026)
by: Sun, Guodong, et al.
Published: (2026)
Complementarity-Free Multi-Contact Modeling and Optimization for Dexterous Manipulation
by: Jin, Wanxin
Published: (2024)
by: Jin, Wanxin
Published: (2024)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
by: Liang, Yujia, et al.
Published: (2025)
by: Liang, Yujia, et al.
Published: (2025)
UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model
by: Li, Zhaowei, et al.
Published: (2024)
by: Li, Zhaowei, et al.
Published: (2024)
Fast Underwater Scene Reconstruction using Multi-View Stereo and Physical Imaging
by: Hu, Shuyi, et al.
Published: (2025)
by: Hu, Shuyi, et al.
Published: (2025)
EMind: A Foundation Model for Multi-task Electromagnetic Signals Understanding
by: Luo, Luqing, et al.
Published: (2025)
by: Luo, Luqing, et al.
Published: (2025)
Multi-modality Anomaly Segmentation on the Road
by: Gao, Heng, et al.
Published: (2025)
by: Gao, Heng, et al.
Published: (2025)
Understanding the Essence: Delving into Annotator Prototype Learning for Multi-Class Annotation Aggregation
by: Chen, Ju, et al.
Published: (2025)
by: Chen, Ju, et al.
Published: (2025)
MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded Scenes
by: Wu, Chenyang, et al.
Published: (2024)
by: Wu, Chenyang, et al.
Published: (2024)
MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval
by: Wu, Xixi, et al.
Published: (2025)
by: Wu, Xixi, et al.
Published: (2025)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
Language as an Anchor: Preserving Relative Visual Geometry for Domain Incremental Learning
by: Geng, Shuyi, et al.
Published: (2025)
by: Geng, Shuyi, et al.
Published: (2025)
Fine-grained Action Analysis: A Multi-modality and Multi-task Dataset of Figure Skating
by: Liu, Sheng-Lan, et al.
Published: (2023)
by: Liu, Sheng-Lan, et al.
Published: (2023)
Distribution-Consistency-Guided Multi-modal Hashing
by: Liu, Jin-Yu, et al.
Published: (2024)
by: Liu, Jin-Yu, et al.
Published: (2024)
MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding
by: Liang, Jing, et al.
Published: (2025)
by: Liang, Jing, et al.
Published: (2025)
3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
by: Li, Zeju, et al.
Published: (2024)
by: Li, Zeju, et al.
Published: (2024)
Part-Whole Relational Fusion Towards Multi-Modal Scene Understanding
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection
by: Zhao, Xiran, et al.
Published: (2026)
by: Zhao, Xiran, et al.
Published: (2026)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
by: Ohi, Masanari, et al.
Published: (2024)
by: Ohi, Masanari, et al.
Published: (2024)
MLVU: Benchmarking Multi-task Long Video Understanding
by: Zhou, Junjie, et al.
Published: (2024)
by: Zhou, Junjie, et al.
Published: (2024)
MTCNET: Multi-task Learning Paradigm for Crowd Count Estimation
by: Kumar, Abhay, et al.
Published: (2019)
by: Kumar, Abhay, et al.
Published: (2019)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
by: Walimbe, Soham, et al.
Published: (2025)
by: Walimbe, Soham, et al.
Published: (2025)
Securing the Floor and Raising the Ceiling: A Merging-based Paradigm for Multi-modal Search Agents
by: Wang, Zhixiang, et al.
Published: (2026)
by: Wang, Zhixiang, et al.
Published: (2026)
Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data
by: Zhu, Xun, et al.
Published: (2025)
by: Zhu, Xun, et al.
Published: (2025)
A Low-Cost, High-Precision Human-Machine Interaction Solution Based on Multi-Coil Wireless Charging Pads
by: Zhang, Bojun
Published: (2025)
by: Zhang, Bojun
Published: (2025)
3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection
by: Ng, Thye Shan, et al.
Published: (2024)
by: Ng, Thye Shan, et al.
Published: (2024)
Online Multi-modal Root Cause Identification in Microservice Systems
by: Zheng, Lecheng, et al.
Published: (2024)
by: Zheng, Lecheng, et al.
Published: (2024)
Similar Items
-
Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey
by: Fu, Ao, et al.
Published: (2024) -
Multi‐modality multiorgan image segmentation using continual learning with enhanced hard attention to the task
by: Ming‐Long Wu, et al.
Published: (2025) -
Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking
by: Zhang, Haonan, et al.
Published: (2025) -
360+x: A Panoptic Multi-modal Scene Understanding Dataset
by: Chen, Hao, et al.
Published: (2024) -
SocialGesture: Delving into Multi-person Gesture Understanding
by: Cao, Xu, et al.
Published: (2025)