ESAM++: Efficient Online 3D Perception on the Edge

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Qin, Aggarwal, Lavisha, Bandyopadhyay, Saptarashmi, Bahirwani, Vikas, Niethammer, Marc, Adeli, Ehsan, Colaco, Andrea
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914612804321280
author Liu, Qin
Aggarwal, Lavisha
Bandyopadhyay, Saptarashmi
Bahirwani, Vikas
Niethammer, Marc
Adeli, Ehsan
Colaco, Andrea
author_facet Liu, Qin
Aggarwal, Lavisha
Bandyopadhyay, Saptarashmi
Bahirwani, Vikas
Niethammer, Marc
Adeli, Ehsan
Colaco, Andrea
contents Online 3D scene perception in real time is essential for robotics, AR/VR, and autonomous systems, particularly in edge computing scenarios where computational resources are limited and privacy is crucial. Recent state-of-the-art methods like EmbodiedSAM (ESAM) demonstrate the promise of online 3D perception by leveraging the Segment Anything Model (SAM) for real-time, fine-grained, and generalized 3D instance segmentation. However, ESAM still relies on a computationally expensive 3D sparse UNet for point cloud feature extraction, which accounts for the majority of the 3D inference time, hindering its practicality on resource-constrained devices. In this paper, we propose ESAM++, a lightweight and scalable alternative for online 3D scene perception tailored to edge devices without GPU acceleration. Our method introduces a 3D Sparse Feature Pyramid Network (SFPN) that efficiently captures multi-scale geometric features from streaming 3D point clouds while significantly reducing computational overhead and model size. We evaluate our approach on four challenging segmentation benchmarks, namely ScanNet, ScanNet200, SceneNN, and 3RScan, demonstrating that our model achieves competitive accuracy with up to 3 times faster inference with a 2 times smaller model size compared to ESAM, enabling practical deployment on edge devices.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29505
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ESAM++: Efficient Online 3D Perception on the Edge
Liu, Qin
Aggarwal, Lavisha
Bandyopadhyay, Saptarashmi
Bahirwani, Vikas
Niethammer, Marc
Adeli, Ehsan
Colaco, Andrea
Computer Vision and Pattern Recognition
Online 3D scene perception in real time is essential for robotics, AR/VR, and autonomous systems, particularly in edge computing scenarios where computational resources are limited and privacy is crucial. Recent state-of-the-art methods like EmbodiedSAM (ESAM) demonstrate the promise of online 3D perception by leveraging the Segment Anything Model (SAM) for real-time, fine-grained, and generalized 3D instance segmentation. However, ESAM still relies on a computationally expensive 3D sparse UNet for point cloud feature extraction, which accounts for the majority of the 3D inference time, hindering its practicality on resource-constrained devices. In this paper, we propose ESAM++, a lightweight and scalable alternative for online 3D scene perception tailored to edge devices without GPU acceleration. Our method introduces a 3D Sparse Feature Pyramid Network (SFPN) that efficiently captures multi-scale geometric features from streaming 3D point clouds while significantly reducing computational overhead and model size. We evaluate our approach on four challenging segmentation benchmarks, namely ScanNet, ScanNet200, SceneNN, and 3RScan, demonstrating that our model achieves competitive accuracy with up to 3 times faster inference with a 2 times smaller model size compared to ESAM, enabling practical deployment on edge devices.
title ESAM++: Efficient Online 3D Perception on the Edge
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.29505