SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kargin, Turhan Can, Jasiński, Wojciech, Pardyl, Adam, Zieliński, Bartosz, Przewięźlikowski, Marcin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond [cls]: Exploring the true potential of Masked Image Modeling representations
by: Przewięźlikowski, Marcin, et al.
Published: (2024)
by: Przewięźlikowski, Marcin, et al.
Published: (2024)
AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale
by: Pardyl, Adam, et al.
Published: (2024)
by: Pardyl, Adam, et al.
Published: (2024)
Beyond Grids: Exploring Elastic Input Sampling for Vision Transformers
by: Pardyl, Adam, et al.
Published: (2023)
by: Pardyl, Adam, et al.
Published: (2023)
Augmentation-aware Self-supervised Learning with Conditioned Projector
by: Przewięźlikowski, Marcin, et al.
Published: (2023)
by: Przewięźlikowski, Marcin, et al.
Published: (2023)
FlySearch: Exploring how vision-language models explore
by: Pardyl, Adam, et al.
Published: (2025)
by: Pardyl, Adam, et al.
Published: (2025)
SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
by: Gong, Ziyang, et al.
Published: (2025)
by: Gong, Ziyang, et al.
Published: (2025)
Parameter-Efficient Interventions for Enhanced Model Merging
by: Osial, Marcin, et al.
Published: (2024)
by: Osial, Marcin, et al.
Published: (2024)
Efficient Multi-Source Knowledge Transfer by Model Merging
by: Osial, Marcin, et al.
Published: (2025)
by: Osial, Marcin, et al.
Published: (2025)
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
by: Ogezi, Michael, et al.
Published: (2025)
by: Ogezi, Michael, et al.
Published: (2025)
Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning
by: Jiang, Haoyi, et al.
Published: (2026)
by: Jiang, Haoyi, et al.
Published: (2026)
SpaGBOL: Spatial-Graph-Based Orientated Localisation
by: Shore, Tavis, et al.
Published: (2024)
by: Shore, Tavis, et al.
Published: (2024)
HyperPlanes: Hypernetwork Approach to Rapid NeRF Adaptation
by: Batorski, Paweł, et al.
Published: (2024)
by: Batorski, Paweł, et al.
Published: (2024)
SpaRTAN: Spatial Reinforcement Token-based Aggregation Network for Visual Recognition
by: Pay, Quan Bi, et al.
Published: (2025)
by: Pay, Quan Bi, et al.
Published: (2025)
Audio-Visual Intelligence in Large Foundation Models
by: Qin, You, et al.
Published: (2026)
by: Qin, You, et al.
Published: (2026)
Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
by: Mirjalili, Vahid, et al.
Published: (2025)
by: Mirjalili, Vahid, et al.
Published: (2025)
SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation
by: Li, Pengna, et al.
Published: (2026)
by: Li, Pengna, et al.
Published: (2026)
SpaCRD: Multimodal Deep Fusion of Histology and Spatial Transcriptomics for Cancer Region Detection
by: Xue, Shuailin, et al.
Published: (2026)
by: Xue, Shuailin, et al.
Published: (2026)
SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation
by: Zhou, Xiaolong, et al.
Published: (2026)
by: Zhou, Xiaolong, et al.
Published: (2026)
SpaRED benchmark: Enhancing Gene Expression Prediction from Histology Images with Spatial Transcriptomics Completion
by: Mejia, Gabriel, et al.
Published: (2024)
by: Mejia, Gabriel, et al.
Published: (2024)
ProtoQuant: Quantization of Prototypical Parts For General and Fine-Grained Image Classification
by: Janusz, Mikołaj, et al.
Published: (2026)
by: Janusz, Mikołaj, et al.
Published: (2026)
Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models
by: Wang, Zengbin, et al.
Published: (2026)
by: Wang, Zengbin, et al.
Published: (2026)
Benchmarking Pathology Foundation Models for Spatial Domain Understanding
by: Zhao, Bokai, et al.
Published: (2026)
by: Zhao, Bokai, et al.
Published: (2026)
PoseCompass: Intelligent Synthetic Pose Selection for Visual Localization
by: Zhou, Yanan, et al.
Published: (2026)
by: Zhou, Yanan, et al.
Published: (2026)
DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
by: Zhang, Ziang, et al.
Published: (2025)
by: Zhang, Ziang, et al.
Published: (2025)
Research on the Spatial Data Intelligent Foundation Model
by: Wang, Shaohua, et al.
Published: (2024)
by: Wang, Shaohua, et al.
Published: (2024)
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
by: Dünkel, Olaf, et al.
Published: (2026)
by: Dünkel, Olaf, et al.
Published: (2026)
Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos
by: Zhao, Zecheng, et al.
Published: (2025)
by: Zhao, Zecheng, et al.
Published: (2025)
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
by: Huang, Xinmiao, et al.
Published: (2025)
by: Huang, Xinmiao, et al.
Published: (2025)
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
by: Wang, Pan, et al.
Published: (2025)
by: Wang, Pan, et al.
Published: (2025)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
by: Yang, Yuchen, et al.
Published: (2026)
by: Yang, Yuchen, et al.
Published: (2026)
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
by: Zhang, Yiming, et al.
Published: (2026)
by: Zhang, Yiming, et al.
Published: (2026)
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
by: Li, Linjie, et al.
Published: (2025)
by: Li, Linjie, et al.
Published: (2025)
Underwater Monocular Metric Depth Estimation: Real-World Benchmarks and Synthetic Fine-Tuning with Vision Foundation Models
by: Cai, Zijie, et al.
Published: (2025)
by: Cai, Zijie, et al.
Published: (2025)
OMENN: One Matrix to Explain Neural Networks
by: Wróbel, Adam, et al.
Published: (2024)
by: Wróbel, Adam, et al.
Published: (2024)
SITE: towards Spatial Intelligence Thorough Evaluation
by: Wang, Wenqi, et al.
Published: (2025)
by: Wang, Wenqi, et al.
Published: (2025)
From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
by: Li, Zongzhao, et al.
Published: (2025)
by: Li, Zongzhao, et al.
Published: (2025)
ForestSim: A Synthetic Benchmark for Intelligent Vehicle Perception in Unstructured Forest Environments
by: Wagle, Pragat, et al.
Published: (2026)
by: Wagle, Pragat, et al.
Published: (2026)
ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models
by: Shen, Qirui, et al.
Published: (2026)
by: Shen, Qirui, et al.
Published: (2026)
PhilEO Bench: Evaluating Geo-Spatial Foundation Models
by: Fibaek, Casper, et al.
Published: (2024)
by: Fibaek, Casper, et al.
Published: (2024)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
Similar Items
-
Beyond [cls]: Exploring the true potential of Masked Image Modeling representations
by: Przewięźlikowski, Marcin, et al.
Published: (2024) -
AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale
by: Pardyl, Adam, et al.
Published: (2024) -
Beyond Grids: Exploring Elastic Input Sampling for Vision Transformers
by: Pardyl, Adam, et al.
Published: (2023) -
Augmentation-aware Self-supervised Learning with Conditioned Projector
by: Przewięźlikowski, Marcin, et al.
Published: (2023) -
FlySearch: Exploring how vision-language models explore
by: Pardyl, Adam, et al.
Published: (2025)