The Bare Necessities: Designing Simple, Effective Open-Vocabulary Scene Graphs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kassab, Christina, Mattamala, Matías, Morin, Sacha, Büchner, Martin, Valada, Abhinav, Paull, Liam, Fallon, Maurice
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929610625646592
author Kassab, Christina
Mattamala, Matías
Morin, Sacha
Büchner, Martin
Valada, Abhinav
Paull, Liam
Fallon, Maurice
author_facet Kassab, Christina
Mattamala, Matías
Morin, Sacha
Büchner, Martin
Valada, Abhinav
Paull, Liam
Fallon, Maurice
contents 3D open-vocabulary scene graph methods are a promising map representation for embodied agents, however many current approaches are computationally expensive. In this paper, we reexamine the critical design choices established in previous works to optimize both efficiency and performance. We propose a general scene graph framework and conduct three studies that focus on image pre-processing, feature fusion, and feature selection. Our findings reveal that commonly used image pre-processing techniques provide minimal performance improvement while tripling computation (on a per object view basis). We also show that averaging feature labels across different views significantly degrades performance. We study alternative feature selection strategies that enhance performance without adding unnecessary computational costs. Based on our findings, we introduce a computationally balanced approach for 3D point cloud segmentation with per-object features. The approach matches state-of-the-art classification accuracy while achieving a threefold reduction in computation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01539
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Bare Necessities: Designing Simple, Effective Open-Vocabulary Scene Graphs
Kassab, Christina
Mattamala, Matías
Morin, Sacha
Büchner, Martin
Valada, Abhinav
Paull, Liam
Fallon, Maurice
Computer Vision and Pattern Recognition
Robotics
3D open-vocabulary scene graph methods are a promising map representation for embodied agents, however many current approaches are computationally expensive. In this paper, we reexamine the critical design choices established in previous works to optimize both efficiency and performance. We propose a general scene graph framework and conduct three studies that focus on image pre-processing, feature fusion, and feature selection. Our findings reveal that commonly used image pre-processing techniques provide minimal performance improvement while tripling computation (on a per object view basis). We also show that averaging feature labels across different views significantly degrades performance. We study alternative feature selection strategies that enhance performance without adding unnecessary computational costs. Based on our findings, we introduce a computationally balanced approach for 3D point cloud segmentation with per-object features. The approach matches state-of-the-art classification accuracy while achieving a threefold reduction in computation.
title The Bare Necessities: Designing Simple, Effective Open-Vocabulary Scene Graphs
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2412.01539