Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miao, Xingyu, Duan, Haoran, Bai, Yang, Shah, Tejal, Song, Jun, Long, Yang, Ranjan, Rajiv, Shao, Ling
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910807264067584
author Miao, Xingyu
Duan, Haoran
Bai, Yang
Shah, Tejal
Song, Jun
Long, Yang
Ranjan, Rajiv
Shao, Ling
author_facet Miao, Xingyu
Duan, Haoran
Bai, Yang
Shah, Tejal
Song, Jun
Long, Yang
Ranjan, Rajiv
Shao, Ling
contents In this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on multi-scale CLIP features and are limited by processing speed and storage requirements, our approach aims to streamline the workflow by directly and effectively distilling dense CLIP features, thereby achieving precise segmentation of 3D scenes using text. To achieve this, we introduce an adapter module and mitigate the noise issue in the dense CLIP feature distillation process through a self-cross-training strategy. Moreover, to enhance the accuracy of segmentation edges, this work presents a low-rank transient query attention mechanism. To ensure the consistency of segmentation for similar colors under different viewpoints, we convert the segmentation task into a classification task through label volume, which significantly improves the consistency of segmentation in color-similar areas. We also propose a simplified text augmentation strategy to alleviate the issue of ambiguity in the correspondence between CLIP features and text. Extensive experimental results show that our method surpasses current state-of-the-art technologies in both training speed and performance. Our code is available on: https://github.com/xingy038/Laser.git.
format Preprint
id arxiv_https___arxiv_org_abs_2501_19084
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields
Miao, Xingyu
Duan, Haoran
Bai, Yang
Shah, Tejal
Song, Jun
Long, Yang
Ranjan, Rajiv
Shao, Ling
Computer Vision and Pattern Recognition
In this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on multi-scale CLIP features and are limited by processing speed and storage requirements, our approach aims to streamline the workflow by directly and effectively distilling dense CLIP features, thereby achieving precise segmentation of 3D scenes using text. To achieve this, we introduce an adapter module and mitigate the noise issue in the dense CLIP feature distillation process through a self-cross-training strategy. Moreover, to enhance the accuracy of segmentation edges, this work presents a low-rank transient query attention mechanism. To ensure the consistency of segmentation for similar colors under different viewpoints, we convert the segmentation task into a classification task through label volume, which significantly improves the consistency of segmentation in color-similar areas. We also propose a simplified text augmentation strategy to alleviate the issue of ambiguity in the correspondence between CLIP features and text. Extensive experimental results show that our method surpasses current state-of-the-art technologies in both training speed and performance. Our code is available on: https://github.com/xingy038/Laser.git.
title Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.19084