Saved in:
Bibliographic Details
Main Authors: Lilova, Valentina, Chakravorty, Toyesh, Bibo, Julian I., Boccaletti, Emma, Li, Brandon, Baxová, Lívia, Snoek, Cees G. M., Salehi, Mohammadreza
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.11574
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914258001854464
author Lilova, Valentina
Chakravorty, Toyesh
Bibo, Julian I.
Boccaletti, Emma
Li, Brandon
Baxová, Lívia
Snoek, Cees G. M.
Salehi, Mohammadreza
author_facet Lilova, Valentina
Chakravorty, Toyesh
Bibo, Julian I.
Boccaletti, Emma
Li, Brandon
Baxová, Lívia
Snoek, Cees G. M.
Salehi, Mohammadreza
contents Benchmarking 3D spatial understanding of foundation models is essential for real-world applications such as robotics and autonomous driving. Existing evaluations often rely on downstream fine-tuning with linear heads or task-specific decoders, making it difficult to isolate the intrinsic 3D reasoning ability of pre-trained encoders. In this work, we introduce a novel benchmark for in-context 3D scene understanding that requires no fine-tuning and directly probes the quality of dense visual features. Building on the Hummingbird framework, which evaluates in-context 2D scene understanding, we extend the setup to the 3D Multi-View ImageNet (MVImgNet) dataset. Given a set of images depicting objects at specific camera angles (keys), we benchmark the performance of segmenting novel views (queries) and report the scores in 4 categories of easy, medium, hard, and extreme based on the key-query view contrast. We benchmark 7 state-of-the-art foundation models and show that DINO-based encoders remain competitive across large viewpoint shifts. Our code is publicly available at https://github.com/ToyeshC/open-hummingbird-3d-eval.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11574
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis
Lilova, Valentina
Chakravorty, Toyesh
Bibo, Julian I.
Boccaletti, Emma
Li, Brandon
Baxová, Lívia
Snoek, Cees G. M.
Salehi, Mohammadreza
Computer Vision and Pattern Recognition
I.4.6
Benchmarking 3D spatial understanding of foundation models is essential for real-world applications such as robotics and autonomous driving. Existing evaluations often rely on downstream fine-tuning with linear heads or task-specific decoders, making it difficult to isolate the intrinsic 3D reasoning ability of pre-trained encoders. In this work, we introduce a novel benchmark for in-context 3D scene understanding that requires no fine-tuning and directly probes the quality of dense visual features. Building on the Hummingbird framework, which evaluates in-context 2D scene understanding, we extend the setup to the 3D Multi-View ImageNet (MVImgNet) dataset. Given a set of images depicting objects at specific camera angles (keys), we benchmark the performance of segmenting novel views (queries) and report the scores in 4 categories of easy, medium, hard, and extreme based on the key-query view contrast. We benchmark 7 state-of-the-art foundation models and show that DINO-based encoders remain competitive across large viewpoint shifts. Our code is publicly available at https://github.com/ToyeshC/open-hummingbird-3d-eval.
title Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis
topic Computer Vision and Pattern Recognition
I.4.6
url https://arxiv.org/abs/2512.11574