Improving Multimodal Distillation for 3D Semantic Segmentation under Domain Shift

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Michele, Björn, Boulch, Alexandre, Puy, Gilles, Vu, Tuan-Hung, Marlet, Renaud, Courty, Nicolas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918213634228224
author Michele, Björn
Boulch, Alexandre
Puy, Gilles
Vu, Tuan-Hung
Marlet, Renaud
Courty, Nicolas
author_facet Michele, Björn
Boulch, Alexandre
Puy, Gilles
Vu, Tuan-Hung
Marlet, Renaud
Courty, Nicolas
contents Semantic segmentation networks trained under full supervision for one type of lidar fail to generalize to unseen lidars without intervention. To reduce the performance gap under domain shifts, a recent trend is to leverage vision foundation models (VFMs) providing robust features across domains. In this work, we conduct an exhaustive study to identify recipes for exploiting VFMs in unsupervised domain adaptation for semantic segmentation of lidar point clouds. Building upon unsupervised image-to-lidar knowledge distillation, our study reveals that: (1) the architecture of the lidar backbone is key to maximize the generalization performance on a target domain; (2) it is possible to pretrain a single backbone once and for all, and use it to address many domain shifts; (3) best results are obtained by keeping the pretrained backbone frozen and training an MLP head for semantic segmentation. The resulting pipeline achieves state-of-the-art results in four widely-recognized and challenging settings. The code will be available at: https://github.com/valeoai/muddos.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17455
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Multimodal Distillation for 3D Semantic Segmentation under Domain Shift
Michele, Björn
Boulch, Alexandre
Puy, Gilles
Vu, Tuan-Hung
Marlet, Renaud
Courty, Nicolas
Computer Vision and Pattern Recognition
Semantic segmentation networks trained under full supervision for one type of lidar fail to generalize to unseen lidars without intervention. To reduce the performance gap under domain shifts, a recent trend is to leverage vision foundation models (VFMs) providing robust features across domains. In this work, we conduct an exhaustive study to identify recipes for exploiting VFMs in unsupervised domain adaptation for semantic segmentation of lidar point clouds. Building upon unsupervised image-to-lidar knowledge distillation, our study reveals that: (1) the architecture of the lidar backbone is key to maximize the generalization performance on a target domain; (2) it is possible to pretrain a single backbone once and for all, and use it to address many domain shifts; (3) best results are obtained by keeping the pretrained backbone frozen and training an MLP head for semantic segmentation. The resulting pipeline achieves state-of-the-art results in four widely-recognized and challenging settings. The code will be available at: https://github.com/valeoai/muddos.
title Improving Multimodal Distillation for 3D Semantic Segmentation under Domain Shift
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.17455