Aligning Text, Images, and 3D Structure Token-by-Token

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sahoo, Aadarsh, Tibrewal, Vansh, Gkioxari, Georgia
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911356669657088
author Sahoo, Aadarsh
Tibrewal, Vansh
Gkioxari, Georgia
author_facet Sahoo, Aadarsh
Tibrewal, Vansh
Gkioxari, Georgia
contents Creating machines capable of understanding the world in 3D is essential in assisting designers that build and edit 3D environments and robots navigating and interacting within a three-dimensional space. Inspired by advances in language and image modeling, we investigate the potential of autoregressive models for a new modality: structured 3D scenes. To this end, we propose a unified LLM framework that aligns language, images, and 3D scenes and provide a detailed ''cookbook'' outlining critical design choices for achieving optimal training and performance addressing key questions related to data representation, modality-specific objectives, and more. We show how to tokenize complex 3D objects to incorporate into our structured 3D scene modality. We evaluate performance across four core 3D tasks -- rendering, recognition, instruction-following, and question-answering -- and four 3D datasets, synthetic and real-world. We show our model's effectiveness on reconstructing complete 3D scenes consisting of complex objects from a single image and on real-world 3D object recognition tasks. Project webpage: https://glab-caltech.github.io/kyvo/
format Preprint
id arxiv_https___arxiv_org_abs_2506_08002
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aligning Text, Images, and 3D Structure Token-by-Token
Sahoo, Aadarsh
Tibrewal, Vansh
Gkioxari, Georgia
Computer Vision and Pattern Recognition
Creating machines capable of understanding the world in 3D is essential in assisting designers that build and edit 3D environments and robots navigating and interacting within a three-dimensional space. Inspired by advances in language and image modeling, we investigate the potential of autoregressive models for a new modality: structured 3D scenes. To this end, we propose a unified LLM framework that aligns language, images, and 3D scenes and provide a detailed ''cookbook'' outlining critical design choices for achieving optimal training and performance addressing key questions related to data representation, modality-specific objectives, and more. We show how to tokenize complex 3D objects to incorporate into our structured 3D scene modality. We evaluate performance across four core 3D tasks -- rendering, recognition, instruction-following, and question-answering -- and four 3D datasets, synthetic and real-world. We show our model's effectiveness on reconstructing complete 3D scenes consisting of complex objects from a single image and on real-world 3D object recognition tasks. Project webpage: https://glab-caltech.github.io/kyvo/
title Aligning Text, Images, and 3D Structure Token-by-Token
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.08002