ELiTe: Efficient Image-to-LiDAR Knowledge Transfer for Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhibo, Yang, Ximing, Zhang, Weizhong, Jin, Cheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913343330058240
author Zhang, Zhibo
Yang, Ximing
Zhang, Weizhong
Jin, Cheng
author_facet Zhang, Zhibo
Yang, Ximing
Zhang, Weizhong
Jin, Cheng
contents Cross-modal knowledge transfer enhances point cloud representation learning in LiDAR semantic segmentation. Despite its potential, the \textit{weak teacher challenge} arises due to repetitive and non-diverse car camera images and sparse, inaccurate ground truth labels. To address this, we propose the Efficient Image-to-LiDAR Knowledge Transfer (ELiTe) paradigm. ELiTe introduces Patch-to-Point Multi-Stage Knowledge Distillation, transferring comprehensive knowledge from the Vision Foundation Model (VFM), extensively trained on diverse open-world images. This enables effective knowledge transfer to a lightweight student model across modalities. ELiTe employs Parameter-Efficient Fine-Tuning to strengthen the VFM teacher and expedite large-scale model training with minimal costs. Additionally, we introduce the Segment Anything Model based Pseudo-Label Generation approach to enhance low-quality image labels, facilitating robust semantic representations. Efficient knowledge transfer in ELiTe yields state-of-the-art results on the SemanticKITTI benchmark, outperforming real-time inference models. Our approach achieves this with significantly fewer parameters, confirming its effectiveness and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2405_04121
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ELiTe: Efficient Image-to-LiDAR Knowledge Transfer for Semantic Segmentation
Zhang, Zhibo
Yang, Ximing
Zhang, Weizhong
Jin, Cheng
Computer Vision and Pattern Recognition
Cross-modal knowledge transfer enhances point cloud representation learning in LiDAR semantic segmentation. Despite its potential, the \textit{weak teacher challenge} arises due to repetitive and non-diverse car camera images and sparse, inaccurate ground truth labels. To address this, we propose the Efficient Image-to-LiDAR Knowledge Transfer (ELiTe) paradigm. ELiTe introduces Patch-to-Point Multi-Stage Knowledge Distillation, transferring comprehensive knowledge from the Vision Foundation Model (VFM), extensively trained on diverse open-world images. This enables effective knowledge transfer to a lightweight student model across modalities. ELiTe employs Parameter-Efficient Fine-Tuning to strengthen the VFM teacher and expedite large-scale model training with minimal costs. Additionally, we introduce the Segment Anything Model based Pseudo-Label Generation approach to enhance low-quality image labels, facilitating robust semantic representations. Efficient knowledge transfer in ELiTe yields state-of-the-art results on the SemanticKITTI benchmark, outperforming real-time inference models. Our approach achieves this with significantly fewer parameters, confirming its effectiveness and efficiency.
title ELiTe: Efficient Image-to-LiDAR Knowledge Transfer for Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.04121