StreamMOS: Streaming Moving Object Segmentation with Multi-View Perception and Dual-Span Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhiheng, Cui, Yubo, Zhong, Jiexi, Fang, Zheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909424133603328
author Li, Zhiheng
Cui, Yubo
Zhong, Jiexi
Fang, Zheng
author_facet Li, Zhiheng
Cui, Yubo
Zhong, Jiexi
Fang, Zheng
contents Moving object segmentation based on LiDAR is a crucial and challenging task for autonomous driving and mobile robotics. Most approaches explore spatio-temporal information from LiDAR sequences to predict moving objects in the current frame. However, they often focus on transferring temporal cues in a single inference and regard every prediction as independent of others. This may cause inconsistent segmentation results for the same object in different frames. To overcome this issue, we propose a streaming network with a memory mechanism, called StreamMOS, to build the association of features and predictions among multiple inferences. Specifically, we utilize a short-term memory to convey historical features, which can be regarded as spatial prior of moving objects and adopted to enhance current inference by temporal fusion. Meanwhile, we build a long-term memory to store previous predictions and exploit them to refine the present forecast at voxel and instance levels through voting. Besides, we present multi-view encoder with cascade projection and asymmetric convolution to extract motion feature of objects in different representations. Extensive experiments validate that our algorithm gets competitive performance on SemanticKITTI and Sipailou Campus datasets. Code will be released at https://github.com/NEU-REAL/StreamMOS.git.
format Preprint
id arxiv_https___arxiv_org_abs_2407_17905
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StreamMOS: Streaming Moving Object Segmentation with Multi-View Perception and Dual-Span Memory
Li, Zhiheng
Cui, Yubo
Zhong, Jiexi
Fang, Zheng
Computer Vision and Pattern Recognition
Robotics
Moving object segmentation based on LiDAR is a crucial and challenging task for autonomous driving and mobile robotics. Most approaches explore spatio-temporal information from LiDAR sequences to predict moving objects in the current frame. However, they often focus on transferring temporal cues in a single inference and regard every prediction as independent of others. This may cause inconsistent segmentation results for the same object in different frames. To overcome this issue, we propose a streaming network with a memory mechanism, called StreamMOS, to build the association of features and predictions among multiple inferences. Specifically, we utilize a short-term memory to convey historical features, which can be regarded as spatial prior of moving objects and adopted to enhance current inference by temporal fusion. Meanwhile, we build a long-term memory to store previous predictions and exploit them to refine the present forecast at voxel and instance levels through voting. Besides, we present multi-view encoder with cascade projection and asymmetric convolution to extract motion feature of objects in different representations. Extensive experiments validate that our algorithm gets competitive performance on SemanticKITTI and Sipailou Campus datasets. Code will be released at https://github.com/NEU-REAL/StreamMOS.git.
title StreamMOS: Streaming Moving Object Segmentation with Multi-View Perception and Dual-Span Memory
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2407.17905