SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Trong-Thuan, Huynh, Viet-Tham, Nguyen, Quang-Thuc, Nguyen, Hoang-Phuc, Bao, Long Le, Minh, Thai Hoang, Anh, Minh Nguyen, Tien, Thang Nguyen, Thuan, Phat Nguyen, Phong, Huy Nguyen, Thai, Bao Huynh, Nguyen, Vinh-Tiep, Nguyen, Duc-Vu, Pham, Phu-Hoa, Le-Hoang, Minh-Huy, Le, Nguyen-Khang, Nguyen, Minh-Chinh, Ho, Minh-Quan, Tran, Ngoc-Long, Le-Hoang, Hien-Long, Tran, Man-Khoi, Tran, Anh-Duong, Nguyen, Kim, Hung, Quan Nguyen, Thanh, Dat Phan, Van, Hoang Tran, Viet, Tien Huynh, Thien, Nhan Nguyen Viet, Vo, Dinh-Khoi, Nguyen, Van-Loc, Le, Trung-Nghia, Nguyen, Tam V., Tran, Minh-Triet
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911102429822976
author Nguyen, Trong-Thuan
Huynh, Viet-Tham
Nguyen, Quang-Thuc
Nguyen, Hoang-Phuc
Bao, Long Le
Minh, Thai Hoang
Anh, Minh Nguyen
Tien, Thang Nguyen
Thuan, Phat Nguyen
Phong, Huy Nguyen
Thai, Bao Huynh
Nguyen, Vinh-Tiep
Nguyen, Duc-Vu
Pham, Phu-Hoa
Le-Hoang, Minh-Huy
Le, Nguyen-Khang
Nguyen, Minh-Chinh
Ho, Minh-Quan
Tran, Ngoc-Long
Le-Hoang, Hien-Long
Tran, Man-Khoi
Tran, Anh-Duong
Nguyen, Kim
Hung, Quan Nguyen
Thanh, Dat Phan
Van, Hoang Tran
Viet, Tien Huynh
Thien, Nhan Nguyen Viet
Vo, Dinh-Khoi
Nguyen, Van-Loc
Le, Trung-Nghia
Nguyen, Tam V.
Tran, Minh-Triet
author_facet Nguyen, Trong-Thuan
Huynh, Viet-Tham
Nguyen, Quang-Thuc
Nguyen, Hoang-Phuc
Bao, Long Le
Minh, Thai Hoang
Anh, Minh Nguyen
Tien, Thang Nguyen
Thuan, Phat Nguyen
Phong, Huy Nguyen
Thai, Bao Huynh
Nguyen, Vinh-Tiep
Nguyen, Duc-Vu
Pham, Phu-Hoa
Le-Hoang, Minh-Huy
Le, Nguyen-Khang
Nguyen, Minh-Chinh
Ho, Minh-Quan
Tran, Ngoc-Long
Le-Hoang, Hien-Long
Tran, Man-Khoi
Tran, Anh-Duong
Nguyen, Kim
Hung, Quan Nguyen
Thanh, Dat Phan
Van, Hoang Tran
Viet, Tien Huynh
Thien, Nhan Nguyen Viet
Vo, Dinh-Khoi
Nguyen, Van-Loc
Le, Trung-Nghia
Nguyen, Tam V.
Tran, Minh-Triet
contents Recent 3D retrieval systems are typically designed for simple, controlled scenarios, such as identifying an object from a cropped image or a brief description. However, real-world scenarios are more complex, often requiring the recognition of an object in a cluttered scene based on a vague, free-form description. To this end, we present ROOMELSA, a new benchmark designed to evaluate a system's ability to interpret natural language. Specifically, ROOMELSA attends to a specific region within a panoramic room image and accurately retrieves the corresponding 3D model from a large database. In addition, ROOMELSA includes over 1,600 apartment scenes, nearly 5,200 rooms, and more than 44,000 targeted queries. Empirically, while coarse object retrieval is largely solved, only one top-performing model consistently ranked the correct match first across nearly all test cases. Notably, a lightweight CLIP-based model also performed well, although it struggled with subtle variations in materials, part structures, and contextual cues, resulting in occasional errors. These findings highlight the importance of tightly integrating visual and language understanding. By bridging the gap between scene-level grounding and fine-grained 3D retrieval, ROOMELSA establishes a new benchmark for advancing robust, real-world 3D recognition systems.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08781
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
Nguyen, Trong-Thuan
Huynh, Viet-Tham
Nguyen, Quang-Thuc
Nguyen, Hoang-Phuc
Bao, Long Le
Minh, Thai Hoang
Anh, Minh Nguyen
Tien, Thang Nguyen
Thuan, Phat Nguyen
Phong, Huy Nguyen
Thai, Bao Huynh
Nguyen, Vinh-Tiep
Nguyen, Duc-Vu
Pham, Phu-Hoa
Le-Hoang, Minh-Huy
Le, Nguyen-Khang
Nguyen, Minh-Chinh
Ho, Minh-Quan
Tran, Ngoc-Long
Le-Hoang, Hien-Long
Tran, Man-Khoi
Tran, Anh-Duong
Nguyen, Kim
Hung, Quan Nguyen
Thanh, Dat Phan
Van, Hoang Tran
Viet, Tien Huynh
Thien, Nhan Nguyen Viet
Vo, Dinh-Khoi
Nguyen, Van-Loc
Le, Trung-Nghia
Nguyen, Tam V.
Tran, Minh-Triet
Computer Vision and Pattern Recognition
Recent 3D retrieval systems are typically designed for simple, controlled scenarios, such as identifying an object from a cropped image or a brief description. However, real-world scenarios are more complex, often requiring the recognition of an object in a cluttered scene based on a vague, free-form description. To this end, we present ROOMELSA, a new benchmark designed to evaluate a system's ability to interpret natural language. Specifically, ROOMELSA attends to a specific region within a panoramic room image and accurately retrieves the corresponding 3D model from a large database. In addition, ROOMELSA includes over 1,600 apartment scenes, nearly 5,200 rooms, and more than 44,000 targeted queries. Empirically, while coarse object retrieval is largely solved, only one top-performing model consistently ranked the correct match first across nearly all test cases. Notably, a lightweight CLIP-based model also performed well, although it struggled with subtle variations in materials, part structures, and contextual cues, resulting in occasional errors. These findings highlight the importance of tightly integrating visual and language understanding. By bridging the gap between scene-level grounding and fine-grained 3D retrieval, ROOMELSA establishes a new benchmark for advancing robust, real-world 3D recognition systems.
title SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.08781