OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Bo-Wen, Cao, Jiao-Long, Zhang, Xuying, Chen, Yuming, Cheng, Ming-Ming, Hou, Qibin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914046015438848
author Yin, Bo-Wen
Cao, Jiao-Long
Zhang, Xuying
Chen, Yuming
Cheng, Ming-Ming
Hou, Qibin
author_facet Yin, Bo-Wen
Cao, Jiao-Long
Zhang, Xuying
Chen, Yuming
Cheng, Ming-Ming
Hou, Qibin
contents Recent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for multiple visual modalities remains unexplored. In this paper, we propose a novel multi-modal learning framework, termed OmniSegmentor. It has two key innovations: 1) Based on ImageNet, we assemble a large-scale dataset for multi-modal pretraining, called ImageNeXt, which contains five popular visual modalities. 2) We provide an efficient pretraining manner to endow the model with the capacity to encode different modality information in the ImageNeXt. For the first time, we introduce a universal multi-modal pretraining framework that consistently amplifies the model's perceptual capabilities across various scenarios, regardless of the arbitrary combination of the involved modalities. Remarkably, our OmniSegmentor achieves new state-of-the-art records on a wide range of multi-modal semantic segmentation datasets, including NYU Depthv2, EventScape, MFNet, DeLiVER, SUNRGBD, and KITTI-360.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15096
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
Yin, Bo-Wen
Cao, Jiao-Long
Zhang, Xuying
Chen, Yuming
Cheng, Ming-Ming
Hou, Qibin
Computer Vision and Pattern Recognition
Recent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for multiple visual modalities remains unexplored. In this paper, we propose a novel multi-modal learning framework, termed OmniSegmentor. It has two key innovations: 1) Based on ImageNet, we assemble a large-scale dataset for multi-modal pretraining, called ImageNeXt, which contains five popular visual modalities. 2) We provide an efficient pretraining manner to endow the model with the capacity to encode different modality information in the ImageNeXt. For the first time, we introduce a universal multi-modal pretraining framework that consistently amplifies the model's perceptual capabilities across various scenarios, regardless of the arbitrary combination of the involved modalities. Remarkably, our OmniSegmentor achieves new state-of-the-art records on a wide range of multi-modal semantic segmentation datasets, including NYU Depthv2, EventScape, MFNet, DeLiVER, SUNRGBD, and KITTI-360.
title OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.15096