MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yu, Huangyue, Jia, Baoxiong, Chen, Yixin, Yang, Yandan, Li, Puhao, Su, Rongpeng, Li, Jiaxin, Li, Qing, Liang, Wei, Zhu, Song-Chun, Liu, Tengyu, Huang, Siyuan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910928063168512
author Yu, Huangyue
Jia, Baoxiong
Chen, Yixin
Yang, Yandan
Li, Puhao
Su, Rongpeng
Li, Jiaxin
Li, Qing
Liang, Wei
Zhu, Song-Chun
Liu, Tengyu
Huang, Siyuan
author_facet Yu, Huangyue
Jia, Baoxiong
Chen, Yixin
Yang, Yandan
Li, Puhao
Su, Rongpeng
Li, Jiaxin
Li, Qing
Liang, Wei
Zhu, Song-Chun
Liu, Tengyu
Huang, Siyuan
contents Embodied AI (EAI) research requires high-quality, diverse 3D scenes to effectively support skill acquisition, sim-to-real transfer, and generalization. Achieving these quality standards, however, necessitates the precise replication of real-world object diversity. Existing datasets demonstrate that this process heavily relies on artist-driven designs, which demand substantial human effort and present significant scalability challenges. To scalably produce realistic and interactive 3D scenes, we first present MetaScenes, a large-scale, simulatable 3D scene dataset constructed from real-world scans, which includes 15366 objects spanning 831 fine-grained categories. Then, we introduce Scan2Sim, a robust multi-modal alignment model, which enables the automated, high-quality replacement of assets, thereby eliminating the reliance on artist-driven designs for scaling 3D scenes. We further propose two benchmarks to evaluate MetaScenes: a detailed scene synthesis task focused on small item layouts for robotic manipulation and a domain transfer task in vision-and-language navigation (VLN) to validate cross-domain transfer. Results confirm MetaScene's potential to enhance EAI by supporting more generalizable agent learning and sim-to-real applications, introducing new possibilities for EAI research. Project website: https://meta-scenes.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02388
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans
Yu, Huangyue
Jia, Baoxiong
Chen, Yixin
Yang, Yandan
Li, Puhao
Su, Rongpeng
Li, Jiaxin
Li, Qing
Liang, Wei
Zhu, Song-Chun
Liu, Tengyu
Huang, Siyuan
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
Embodied AI (EAI) research requires high-quality, diverse 3D scenes to effectively support skill acquisition, sim-to-real transfer, and generalization. Achieving these quality standards, however, necessitates the precise replication of real-world object diversity. Existing datasets demonstrate that this process heavily relies on artist-driven designs, which demand substantial human effort and present significant scalability challenges. To scalably produce realistic and interactive 3D scenes, we first present MetaScenes, a large-scale, simulatable 3D scene dataset constructed from real-world scans, which includes 15366 objects spanning 831 fine-grained categories. Then, we introduce Scan2Sim, a robust multi-modal alignment model, which enables the automated, high-quality replacement of assets, thereby eliminating the reliance on artist-driven designs for scaling 3D scenes. We further propose two benchmarks to evaluate MetaScenes: a detailed scene synthesis task focused on small item layouts for robotic manipulation and a domain transfer task in vision-and-language navigation (VLN) to validate cross-domain transfer. Results confirm MetaScene's potential to enhance EAI by supporting more generalizable agent learning and sim-to-real applications, introducing new possibilities for EAI research. Project website: https://meta-scenes.github.io/.
title MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2505.02388