Scaling LLaNA: Advancing NeRF-Language Understanding Through Large-Scale Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Amaduzzi, Andrea, Ramirez, Pierluigi Zama, Lisanti, Giuseppe, Salti, Samuele, Di Stefano, Luigi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912335038251008
author Amaduzzi, Andrea
Ramirez, Pierluigi Zama
Lisanti, Giuseppe
Salti, Samuele
Di Stefano, Luigi
author_facet Amaduzzi, Andrea
Ramirez, Pierluigi Zama
Lisanti, Giuseppe
Salti, Samuele
Di Stefano, Luigi
contents Recent advances in Multimodal Large Language Models (MLLMs) have shown remarkable capabilities in understanding both images and 3D data, yet these modalities face inherent limitations in comprehensively representing object geometry and appearance. Neural Radiance Fields (NeRFs) have emerged as a promising alternative, encoding both geometric and photorealistic properties within the weights of a simple Multi-Layer Perceptron (MLP). This work investigates the feasibility and effectiveness of ingesting NeRFs into an MLLM. We introduce LLaNA, the first MLLM able to perform new tasks such as NeRF captioning and Q\&A, by directly processing the weights of a NeRF's MLP. Notably, LLaNA is able to extract information about the represented objects without the need to render images or materialize 3D data structures. In addition, we build the first large-scale NeRF-language dataset, composed by more than 300K NeRFs trained on ShapeNet and Objaverse, with paired textual annotations that enable various NeRF-language tasks. Based on this dataset, we develop a benchmark to evaluate the NeRF understanding capability of our method. Results show that directly processing NeRF weights leads to better performance on NeRF-Language tasks compared to approaches that rely on either 2D or 3D representations derived from NeRFs.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13995
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling LLaNA: Advancing NeRF-Language Understanding Through Large-Scale Training
Amaduzzi, Andrea
Ramirez, Pierluigi Zama
Lisanti, Giuseppe
Salti, Samuele
Di Stefano, Luigi
Computer Vision and Pattern Recognition
Recent advances in Multimodal Large Language Models (MLLMs) have shown remarkable capabilities in understanding both images and 3D data, yet these modalities face inherent limitations in comprehensively representing object geometry and appearance. Neural Radiance Fields (NeRFs) have emerged as a promising alternative, encoding both geometric and photorealistic properties within the weights of a simple Multi-Layer Perceptron (MLP). This work investigates the feasibility and effectiveness of ingesting NeRFs into an MLLM. We introduce LLaNA, the first MLLM able to perform new tasks such as NeRF captioning and Q\&A, by directly processing the weights of a NeRF's MLP. Notably, LLaNA is able to extract information about the represented objects without the need to render images or materialize 3D data structures. In addition, we build the first large-scale NeRF-language dataset, composed by more than 300K NeRFs trained on ShapeNet and Objaverse, with paired textual annotations that enable various NeRF-language tasks. Based on this dataset, we develop a benchmark to evaluate the NeRF understanding capability of our method. Results show that directly processing NeRF weights leads to better performance on NeRF-Language tasks compared to approaches that rely on either 2D or 3D representations derived from NeRFs.
title Scaling LLaNA: Advancing NeRF-Language Understanding Through Large-Scale Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.13995