Does Spatial Cognition Emerge in Frontier Models?
Fuente:
arXiv
Saved in:
| Main Authors: | Ramakrishnan, Santhosh Kumar, Wijmans, Erik, Kraehenbuehl, Philipp, Koltun, Vladlen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do multimodal models imagine electric sheep?
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
Cut Your Losses in Large-Vocabulary Language Models
by: Wijmans, Erik, et al.
Published: (2024)
by: Wijmans, Erik, et al.
Published: (2024)
CoMotion: Concurrent Multi-person 3D Motion
by: Newell, Alejandro, et al.
Published: (2025)
by: Newell, Alejandro, et al.
Published: (2025)
DSFEC: Efficient and Deployable Deep Radar Object Detection
by: Dandugula, Gayathri, et al.
Published: (2024)
by: Dandugula, Gayathri, et al.
Published: (2024)
Breakdance Video classification in the age of Generative AI
by: Dhar, Sauptik, et al.
Published: (2025)
by: Dhar, Sauptik, et al.
Published: (2025)
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
by: Li, Zhiqi, et al.
Published: (2025)
by: Li, Zhiqi, et al.
Published: (2025)
Robust Autonomy Emerges from Self-Play
by: Cusumano-Towner, Marco, et al.
Published: (2025)
by: Cusumano-Towner, Marco, et al.
Published: (2025)
Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
by: Li, Yihao, et al.
Published: (2025)
by: Li, Yihao, et al.
Published: (2025)
Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
by: Bochkovskii, Aleksei, et al.
Published: (2024)
by: Bochkovskii, Aleksei, et al.
Published: (2024)
Doe-1: Closed-Loop Autonomous Driving with Large World Model
by: Zheng, Wenzhao, et al.
Published: (2024)
by: Zheng, Wenzhao, et al.
Published: (2024)
Research on the Spatial Data Intelligent Foundation Model
by: Wang, Shaohua, et al.
Published: (2024)
by: Wang, Shaohua, et al.
Published: (2024)
Spatial Transformer Network YOLO Model for Agricultural Object Detection
by: Zambre, Yash, et al.
Published: (2024)
by: Zambre, Yash, et al.
Published: (2024)
End-to-end Autonomous Driving: Challenges and Frontiers
by: Chen, Li, et al.
Published: (2023)
by: Chen, Li, et al.
Published: (2023)
When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models
by: Al-Hamadani, Samer
Published: (2025)
by: Al-Hamadani, Samer
Published: (2025)
AmaraSpatial-10K: A Spatially and Semantically Aligned 3D Dataset for Spatial Computing and Embodied AI
by: Salehi, Mohammad Sadegh, et al.
Published: (2026)
by: Salehi, Mohammad Sadegh, et al.
Published: (2026)
VIPaint: Image Inpainting with Pre-Trained Diffusion Models via Variational Inference
by: Agarwal, Sakshi, et al.
Published: (2024)
by: Agarwal, Sakshi, et al.
Published: (2024)
Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks
by: Sharma, Arun
Published: (2026)
by: Sharma, Arun
Published: (2026)
Does a Neural Network Really Encode Symbolic Concepts?
by: Li, Mingjie, et al.
Published: (2023)
by: Li, Mingjie, et al.
Published: (2023)
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
by: Ramachandran, Rahul, et al.
Published: (2025)
by: Ramachandran, Rahul, et al.
Published: (2025)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
by: Zhang, Zhengshen, et al.
Published: (2025)
by: Zhang, Zhengshen, et al.
Published: (2025)
Does FLUX Already Know How to Perform Physically Plausible Image Composition?
by: Lu, Shilin, et al.
Published: (2025)
by: Lu, Shilin, et al.
Published: (2025)
Sparks of Artificial General Intelligence(AGI) in Semiconductor Material Science: Early Explorations into the Next Frontier of Generative AI-Assisted Electron Micrograph Analysis
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Spatial-Relation Data
by: Yamamoto, Takaki, et al.
Published: (2026)
by: Yamamoto, Takaki, et al.
Published: (2026)
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
by: Mayilvahanan, Prasanna, et al.
Published: (2023)
by: Mayilvahanan, Prasanna, et al.
Published: (2023)
When Better Eyes Lead to Blindness: A Diagnostic Study of the Information Bottleneck in CNN-LSTM Image Captioning Models
by: Gupta, Hitesh Kumar
Published: (2025)
by: Gupta, Hitesh Kumar
Published: (2025)
An Interpretable Implicit-Based Approach for Modeling Local Spatial Effects: A Case Study of Global Gross Primary Productivity
by: Du, Siqi, et al.
Published: (2025)
by: Du, Siqi, et al.
Published: (2025)
On Spectral Properties of Gradient-based Explanation Methods
by: Mehrpanah, Amir, et al.
Published: (2025)
by: Mehrpanah, Amir, et al.
Published: (2025)
Analysis of Hyperparameter Optimization Effects on Lightweight Deep Models for Real-Time Image Classification
by: Rakesh, Vineet Kumar, et al.
Published: (2025)
by: Rakesh, Vineet Kumar, et al.
Published: (2025)
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
by: Rutar, Danaja, et al.
Published: (2025)
by: Rutar, Danaja, et al.
Published: (2025)
Autonomous AI Surveillance: Multimodal Deep Learning for Cognitive and Behavioral Monitoring
by: Hamza, Ameer, et al.
Published: (2025)
by: Hamza, Ameer, et al.
Published: (2025)
Frequency-Aware Model Parameter Explorer: A new attribution method for improving explainability
by: Yavari, Ali, et al.
Published: (2025)
by: Yavari, Ali, et al.
Published: (2025)
SENSE: Self-Supervised Neural Embeddings for Spatial Ensembles
by: Gadirov, Hamid, et al.
Published: (2025)
by: Gadirov, Hamid, et al.
Published: (2025)
Longitudinal Flow Matching for Trajectory Modeling
by: Islam, Mohammad Mohaiminul, et al.
Published: (2025)
by: Islam, Mohammad Mohaiminul, et al.
Published: (2025)
RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and Exploration
by: Alama, Omar, et al.
Published: (2025)
by: Alama, Omar, et al.
Published: (2025)
A Markovian View of Iterative-Feedback Loops in Image Generative Models: Neural Resonance and Model Collapse
by: Vats, Vibhas Kumar, et al.
Published: (2026)
by: Vats, Vibhas Kumar, et al.
Published: (2026)
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy
by: Pattnayak, Priyaranjan, et al.
Published: (2024)
by: Pattnayak, Priyaranjan, et al.
Published: (2024)
How Well Does GPT-4V(ision) Adapt to Distribution Shifts? A Preliminary Investigation
by: Han, Zhongyi, et al.
Published: (2023)
by: Han, Zhongyi, et al.
Published: (2023)
Personalized Federated Training of Diffusion Models with Privacy Guarantees
by: Patel, Kumar Kshitij, et al.
Published: (2025)
by: Patel, Kumar Kshitij, et al.
Published: (2025)
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
by: Zhao, Yaqi, et al.
Published: (2024)
by: Zhao, Yaqi, et al.
Published: (2024)
Similar Items
-
Do multimodal models imagine electric sheep?
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026) -
Cut Your Losses in Large-Vocabulary Language Models
by: Wijmans, Erik, et al.
Published: (2024) -
CoMotion: Concurrent Multi-person 3D Motion
by: Newell, Alejandro, et al.
Published: (2025) -
DSFEC: Efficient and Deployable Deep Radar Object Detection
by: Dandugula, Gayathri, et al.
Published: (2024) -
Breakdance Video classification in the age of Generative AI
by: Dhar, Sauptik, et al.
Published: (2025)