OpenLex3D: A Tiered Evaluation Benchmark for Open-Vocabulary 3D Scene Representations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kassab, Christina, Morin, Sacha, Büchner, Martin, Mattamala, Matías, Gupta, Kumaraditya, Valada, Abhinav, Paull, Liam, Fallon, Maurice
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911207891402752
author Kassab, Christina
Morin, Sacha
Büchner, Martin
Mattamala, Matías
Gupta, Kumaraditya
Valada, Abhinav
Paull, Liam
Fallon, Maurice
author_facet Kassab, Christina
Morin, Sacha
Büchner, Martin
Mattamala, Matías
Gupta, Kumaraditya
Valada, Abhinav
Paull, Liam
Fallon, Maurice
contents 3D scene understanding has been transformed by open-vocabulary language models that enable interaction via natural language. However, at present the evaluation of these representations is limited to datasets with closed-set semantics that do not capture the richness of language. This work presents OpenLex3D, a dedicated benchmark for evaluating 3D open-vocabulary scene representations. OpenLex3D provides entirely new label annotations for scenes from Replica, ScanNet++, and HM3D, which capture real-world linguistic variability by introducing synonymical object categories and additional nuanced descriptions. Our label sets provide 13 times more labels per scene than the original datasets. By introducing an open-set 3D semantic segmentation task and an object retrieval task, we evaluate various existing 3D open-vocabulary methods on OpenLex3D, showcasing failure cases, and avenues for improvement. Our experiments provide insights on feature precision, segmentation, and downstream capabilities. The benchmark is publicly available at: https://openlex3d.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19764
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OpenLex3D: A Tiered Evaluation Benchmark for Open-Vocabulary 3D Scene Representations
Kassab, Christina
Morin, Sacha
Büchner, Martin
Mattamala, Matías
Gupta, Kumaraditya
Valada, Abhinav
Paull, Liam
Fallon, Maurice
Computer Vision and Pattern Recognition
Robotics
3D scene understanding has been transformed by open-vocabulary language models that enable interaction via natural language. However, at present the evaluation of these representations is limited to datasets with closed-set semantics that do not capture the richness of language. This work presents OpenLex3D, a dedicated benchmark for evaluating 3D open-vocabulary scene representations. OpenLex3D provides entirely new label annotations for scenes from Replica, ScanNet++, and HM3D, which capture real-world linguistic variability by introducing synonymical object categories and additional nuanced descriptions. Our label sets provide 13 times more labels per scene than the original datasets. By introducing an open-set 3D semantic segmentation task and an object retrieval task, we evaluate various existing 3D open-vocabulary methods on OpenLex3D, showcasing failure cases, and avenues for improvement. Our experiments provide insights on feature precision, segmentation, and downstream capabilities. The benchmark is publicly available at: https://openlex3d.github.io/.
title OpenLex3D: A Tiered Evaluation Benchmark for Open-Vocabulary 3D Scene Representations
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2503.19764