Large Vision Models Can Solve Mental Rotation Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Mason, Sebastian Ray, Gjølbye, Anders, Højbjerg, Phillip Chavarria, Tětková, Lenka, Hansen, Lars Kai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robustness of Visual Explanations to Common Data Augmentation
by: Tětková, Lenka, et al.
Published: (2023)
by: Tětková, Lenka, et al.
Published: (2023)
Challenges in explaining deep learning models for data with biological variation
by: Tětková, Lenka, et al.
Published: (2024)
by: Tětková, Lenka, et al.
Published: (2024)
From Colors to Classes: Emergence of Concepts in Vision Transformers
by: Dorszewski, Teresa, et al.
Published: (2025)
by: Dorszewski, Teresa, et al.
Published: (2025)
Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven' Matrices
by: Deng, Hokin
Published: (2025)
by: Deng, Hokin
Published: (2025)
GeoBiked: A Dataset with Geometric Features and Automated Labeling Techniques to Enable Deep Generative Models in Engineering Design
by: Mueller, Phillip, et al.
Published: (2024)
by: Mueller, Phillip, et al.
Published: (2024)
Uncovering Bias in Large Vision-Language Models with Counterfactuals
by: Howard, Phillip, et al.
Published: (2024)
by: Howard, Phillip, et al.
Published: (2024)
Can Vision-Language Models Solve Visual Math Equations?
by: Choudhury, Monjoy Narayan, et al.
Published: (2025)
by: Choudhury, Monjoy Narayan, et al.
Published: (2025)
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
by: Yuan, Fan, et al.
Published: (2025)
by: Yuan, Fan, et al.
Published: (2025)
Cross-Cultural Value Awareness in Large Vision-Language Models
by: Howard, Phillip, et al.
Published: (2026)
by: Howard, Phillip, et al.
Published: (2026)
AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?
by: Bao, Han, et al.
Published: (2024)
by: Bao, Han, et al.
Published: (2024)
Using Vision Language Foundation Models to Generate Plant Simulation Configurations via In-Context Learning
by: Yun, Heesup, et al.
Published: (2026)
by: Yun, Heesup, et al.
Published: (2026)
CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences
by: Lin, Fangzhou, et al.
Published: (2026)
by: Lin, Fangzhou, et al.
Published: (2026)
Vision-Language Models under Cultural and Inclusive Considerations
by: Karamolegkou, Antonia, et al.
Published: (2024)
by: Karamolegkou, Antonia, et al.
Published: (2024)
Physical Prompt Injection Attacks on Large Vision-Language Models
by: Ling, Chen, et al.
Published: (2026)
by: Ling, Chen, et al.
Published: (2026)
Bias Detection and Rotation-Robustness Mitigation in Vision-Language Models and Generative Image Models
by: Mithila, Tarannum
Published: (2026)
by: Mithila, Tarannum
Published: (2026)
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
by: Kang, Beomseok, et al.
Published: (2026)
by: Kang, Beomseok, et al.
Published: (2026)
MVP-Bench: Can Large Vision--Language Models Conduct Multi-level Visual Perception Like Humans?
by: Li, Guanzhen, et al.
Published: (2024)
by: Li, Guanzhen, et al.
Published: (2024)
MeshFleet: Filtered and Annotated 3D Vehicle Dataset for Domain Specific Generative Modeling
by: Boborzi, Damian, et al.
Published: (2025)
by: Boborzi, Damian, et al.
Published: (2025)
Relaxed Rotational Equivariance via $G$-Biases in Vision
by: Wu, Zhiqiang, et al.
Published: (2024)
by: Wu, Zhiqiang, et al.
Published: (2024)
Learning Image Priors through Patch-based Diffusion Models for Solving Inverse Problems
by: Hu, Jason, et al.
Published: (2024)
by: Hu, Jason, et al.
Published: (2024)
A Survey on Vision Autoregressive Model
by: Jiang, Kai, et al.
Published: (2024)
by: Jiang, Kai, et al.
Published: (2024)
Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks
by: Dorszewski, Teresa, et al.
Published: (2024)
by: Dorszewski, Teresa, et al.
Published: (2024)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
Can Vision-Language Models Understand Construction Workers? An Exploratory Study
by: Bui, Hieu, et al.
Published: (2026)
by: Bui, Hieu, et al.
Published: (2026)
Assessing Color Vision Test in Large Vision-language Models
by: Ye, Hongfei, et al.
Published: (2025)
by: Ye, Hongfei, et al.
Published: (2025)
MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
by: Li, Chenxi, et al.
Published: (2025)
by: Li, Chenxi, et al.
Published: (2025)
Solving Video Inverse Problems Using Image Diffusion Models
by: Kwon, Taesung, et al.
Published: (2024)
by: Kwon, Taesung, et al.
Published: (2024)
Can Vision-Language Models Solve the Shell Game?
by: Liu, Tiedong, et al.
Published: (2026)
by: Liu, Tiedong, et al.
Published: (2026)
SocialCounterfactuals: Probing and Mitigating Intersectional Social Biases in Vision-Language Models with Counterfactual Examples
by: Howard, Phillip, et al.
Published: (2023)
by: Howard, Phillip, et al.
Published: (2023)
How Lightweight Can A Vision Transformer Be
by: Tan, Jen Hong
Published: (2024)
by: Tan, Jen Hong
Published: (2024)
LayerShuffle: Enhancing Robustness in Vision Transformers by Randomizing Layer Execution Order
by: Freiberger, Matthias, et al.
Published: (2024)
by: Freiberger, Matthias, et al.
Published: (2024)
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
by: Lee, Phillip Y., et al.
Published: (2025)
by: Lee, Phillip Y., et al.
Published: (2025)
Can Large Vision-Language Models Detect Images Copyright Infringement from GenAI?
by: Xu, Qipan, et al.
Published: (2025)
by: Xu, Qipan, et al.
Published: (2025)
Geo-LLaVA: A Large Multi-Modal Model for Solving Geometry Math Problems with Meta In-Context Learning
by: Xu, Shihao, et al.
Published: (2024)
by: Xu, Shihao, et al.
Published: (2024)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
by: Kundu, Souvik, et al.
Published: (2025)
by: Kundu, Souvik, et al.
Published: (2025)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Trustworthy Large Models in Vision: A Survey
by: Guo, Ziyan, et al.
Published: (2023)
by: Guo, Ziyan, et al.
Published: (2023)
Can Vision Models Truly Forget? Mirage: Representation-Level Certification of Visual Unlearning
by: Yu, Zhenyu, et al.
Published: (2026)
by: Yu, Zhenyu, et al.
Published: (2026)
GameVerse: Can Vision-Language Models Learn from Video-based Reflection?
by: Zhang, Kuan, et al.
Published: (2026)
by: Zhang, Kuan, et al.
Published: (2026)
Warped Diffusion: Solving Video Inverse Problems with Image Diffusion Models
by: Daras, Giannis, et al.
Published: (2024)
by: Daras, Giannis, et al.
Published: (2024)
Similar Items
-
Robustness of Visual Explanations to Common Data Augmentation
by: Tětková, Lenka, et al.
Published: (2023) -
Challenges in explaining deep learning models for data with biological variation
by: Tětková, Lenka, et al.
Published: (2024) -
From Colors to Classes: Emergence of Concepts in Vision Transformers
by: Dorszewski, Teresa, et al.
Published: (2025) -
Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven' Matrices
by: Deng, Hokin
Published: (2025) -
GeoBiked: A Dataset with Geometric Features and Automated Labeling Techniques to Enable Deep Generative Models in Engineering Design
by: Mueller, Phillip, et al.
Published: (2024)