Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.11574 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914258001854464 |
|---|---|
| author | Lilova, Valentina Chakravorty, Toyesh Bibo, Julian I. Boccaletti, Emma Li, Brandon Baxová, Lívia Snoek, Cees G. M. Salehi, Mohammadreza |
| author_facet | Lilova, Valentina Chakravorty, Toyesh Bibo, Julian I. Boccaletti, Emma Li, Brandon Baxová, Lívia Snoek, Cees G. M. Salehi, Mohammadreza |
| contents | Benchmarking 3D spatial understanding of foundation models is essential for real-world applications such as robotics and autonomous driving. Existing evaluations often rely on downstream fine-tuning with linear heads or task-specific decoders, making it difficult to isolate the intrinsic 3D reasoning ability of pre-trained encoders. In this work, we introduce a novel benchmark for in-context 3D scene understanding that requires no fine-tuning and directly probes the quality of dense visual features. Building on the Hummingbird framework, which evaluates in-context 2D scene understanding, we extend the setup to the 3D Multi-View ImageNet (MVImgNet) dataset. Given a set of images depicting objects at specific camera angles (keys), we benchmark the performance of segmenting novel views (queries) and report the scores in 4 categories of easy, medium, hard, and extreme based on the key-query view contrast. We benchmark 7 state-of-the-art foundation models and show that DINO-based encoders remain competitive across large viewpoint shifts. Our code is publicly available at https://github.com/ToyeshC/open-hummingbird-3d-eval. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_11574 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis Lilova, Valentina Chakravorty, Toyesh Bibo, Julian I. Boccaletti, Emma Li, Brandon Baxová, Lívia Snoek, Cees G. M. Salehi, Mohammadreza Computer Vision and Pattern Recognition I.4.6 Benchmarking 3D spatial understanding of foundation models is essential for real-world applications such as robotics and autonomous driving. Existing evaluations often rely on downstream fine-tuning with linear heads or task-specific decoders, making it difficult to isolate the intrinsic 3D reasoning ability of pre-trained encoders. In this work, we introduce a novel benchmark for in-context 3D scene understanding that requires no fine-tuning and directly probes the quality of dense visual features. Building on the Hummingbird framework, which evaluates in-context 2D scene understanding, we extend the setup to the 3D Multi-View ImageNet (MVImgNet) dataset. Given a set of images depicting objects at specific camera angles (keys), we benchmark the performance of segmenting novel views (queries) and report the scores in 4 categories of easy, medium, hard, and extreme based on the key-query view contrast. We benchmark 7 state-of-the-art foundation models and show that DINO-based encoders remain competitive across large viewpoint shifts. Our code is publicly available at https://github.com/ToyeshC/open-hummingbird-3d-eval. |
| title | Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis |
| topic | Computer Vision and Pattern Recognition I.4.6 |
| url | https://arxiv.org/abs/2512.11574 |