Anyview: Generalizable Indoor 3D Object Detection with Variable Frames

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Zhenyu, Xu, Xiuwei, Wang, Ziwei, Xia, Chong, Zhao, Linqing, Lu, Jiwen, Yan, Haibin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916820146978816
author Wu, Zhenyu
Xu, Xiuwei
Wang, Ziwei
Xia, Chong
Zhao, Linqing
Lu, Jiwen
Yan, Haibin
author_facet Wu, Zhenyu
Xu, Xiuwei
Wang, Ziwei
Xia, Chong
Zhao, Linqing
Lu, Jiwen
Yan, Haibin
contents In this paper, we propose a novel network framework for indoor 3D object detection to handle variable input frame numbers in practical scenarios. Existing methods only consider fixed frames of input data for a single detector, such as monocular RGB-D images or point clouds reconstructed from dense multi-view RGB-D images. While in practical application scenes such as robot navigation and manipulation, the raw input to the 3D detectors is the RGB-D images with variable frame numbers instead of the reconstructed scene point cloud. However, the previous approaches can only handle fixed frame input data and have poor performance with variable frame input. In order to facilitate 3D object detection methods suitable for practical tasks, we present a novel 3D detection framework named AnyView for our practical applications, which generalizes well across different numbers of input frames with a single model. To be specific, we propose a geometric learner to mine the local geometric features of each input RGB-D image frame and implement local-global feature interaction through a designed spatial mixture module. Meanwhile, we further utilize a dynamic token strategy to adaptively adjust the number of extracted features for each frame, which ensures consistent global feature density and further enhances the generalization after fusion. Extensive experiments on the ScanNet dataset show our method achieves both great generalizability and high detection accuracy with a simple and clean architecture containing a similar amount of parameters with the baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2310_05346
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Anyview: Generalizable Indoor 3D Object Detection with Variable Frames
Wu, Zhenyu
Xu, Xiuwei
Wang, Ziwei
Xia, Chong
Zhao, Linqing
Lu, Jiwen
Yan, Haibin
Computer Vision and Pattern Recognition
Robotics
In this paper, we propose a novel network framework for indoor 3D object detection to handle variable input frame numbers in practical scenarios. Existing methods only consider fixed frames of input data for a single detector, such as monocular RGB-D images or point clouds reconstructed from dense multi-view RGB-D images. While in practical application scenes such as robot navigation and manipulation, the raw input to the 3D detectors is the RGB-D images with variable frame numbers instead of the reconstructed scene point cloud. However, the previous approaches can only handle fixed frame input data and have poor performance with variable frame input. In order to facilitate 3D object detection methods suitable for practical tasks, we present a novel 3D detection framework named AnyView for our practical applications, which generalizes well across different numbers of input frames with a single model. To be specific, we propose a geometric learner to mine the local geometric features of each input RGB-D image frame and implement local-global feature interaction through a designed spatial mixture module. Meanwhile, we further utilize a dynamic token strategy to adaptively adjust the number of extracted features for each frame, which ensures consistent global feature density and further enhances the generalization after fusion. Extensive experiments on the ScanNet dataset show our method achieves both great generalizability and high detection accuracy with a simple and clean architecture containing a similar amount of parameters with the baselines.
title Anyview: Generalizable Indoor 3D Object Detection with Variable Frames
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2310.05346