Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Zhongbin, Yang, Zhen, Li, Yushan, Zhang, Xinyue, Gao, Wenyu, Wang, Jiacheng, Li, Chengzhi, Liu, Xiangrui, Jian, Ping |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Flatlands: Unlocking Spatial Intelligence by Decoupling 3D Reasoning from Numerical Regression
by: Guo, Zhongbin, et al.
Published: (2025)
by: Guo, Zhongbin, et al.
Published: (2025)
LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency
by: Guo, Zhongbin, et al.
Published: (2025)
by: Guo, Zhongbin, et al.
Published: (2025)
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models
by: Guo, Zhongbin, et al.
Published: (2025)
by: Guo, Zhongbin, et al.
Published: (2025)
How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
by: Yang, Zhen, et al.
Published: (2026)
by: Yang, Zhen, et al.
Published: (2026)
The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
by: Deng, Yonghong, et al.
Published: (2026)
by: Deng, Yonghong, et al.
Published: (2026)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
by: Li, Chengzhi, et al.
Published: (2025)
by: Li, Chengzhi, et al.
Published: (2025)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Pixels, Patterns, but No Poetry: To See The World like Humans
by: Gao, Hongcheng, et al.
Published: (2025)
by: Gao, Hongcheng, et al.
Published: (2025)
Discovering Intrinsic Spatial-Temporal Logic Rules to Explain Human Actions
by: Cao, Chengzhi, et al.
Published: (2023)
by: Cao, Chengzhi, et al.
Published: (2023)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
by: Lai, Zhengzhao, et al.
Published: (2025)
by: Lai, Zhengzhao, et al.
Published: (2025)
Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
by: Shen, Yixuan, et al.
Published: (2026)
by: Shen, Yixuan, et al.
Published: (2026)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT
by: Wang, Jiacheng, et al.
Published: (2026)
by: Wang, Jiacheng, et al.
Published: (2026)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025)
by: Feng, Zhiyuan, et al.
Published: (2025)
From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs
by: Zhang, Le, et al.
Published: (2026)
by: Zhang, Le, et al.
Published: (2026)
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
START: Spatial and Textual Learning for Chart Understanding
by: Liu, Zhuoming, et al.
Published: (2025)
by: Liu, Zhuoming, et al.
Published: (2025)
SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation
by: Zhou, Xiaolong, et al.
Published: (2026)
by: Zhou, Xiaolong, et al.
Published: (2026)
Motion Generation from Fine-grained Textual Descriptions
by: Li, Kunhang, et al.
Published: (2024)
by: Li, Kunhang, et al.
Published: (2024)
Advancing Textual Prompt Learning with Anchored Attributes
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence
by: Chen, Jiabin, et al.
Published: (2025)
by: Chen, Jiabin, et al.
Published: (2025)
Embedding Textual Information in Images Using Quinary Pixel Combinations
by: Kandala, A V Uday Kiran
Published: (2026)
by: Kandala, A V Uday Kiran
Published: (2026)
Pixelis: Reasoning in Pixels, from Seeing to Acting
by: Zhou, Yunpeng
Published: (2026)
by: Zhou, Yunpeng
Published: (2026)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
by: Yang, Yuchen, et al.
Published: (2026)
by: Yang, Yuchen, et al.
Published: (2026)
Seeing without Pixels: Perception from Camera Trajectories
by: Xue, Zihui, et al.
Published: (2025)
by: Xue, Zihui, et al.
Published: (2025)
Unsupervised Spatial-Temporal Feature Enrichment and Fidelity Preservation Network for Skeleton based Action Recognition
by: Li, Chuankun, et al.
Published: (2024)
by: Li, Chuankun, et al.
Published: (2024)
See, Remember, Explore: A Benchmark and Baselines for Streaming Spatial Reasoning
by: Wei, Yuxi, et al.
Published: (2026)
by: Wei, Yuxi, et al.
Published: (2026)
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
by: Wang, Pan, et al.
Published: (2025)
by: Wang, Pan, et al.
Published: (2025)
Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach
by: Liu, Feiyang, et al.
Published: (2024)
by: Liu, Feiyang, et al.
Published: (2024)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
by: Hua, Jiacheng, et al.
Published: (2026)
by: Hua, Jiacheng, et al.
Published: (2026)
Can Textual Semantics Mitigate Sounding Object Segmentation Preference?
by: Wang, Yaoting, et al.
Published: (2024)
by: Wang, Yaoting, et al.
Published: (2024)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
by: Li, Xiujun, et al.
Published: (2023)
by: Li, Xiujun, et al.
Published: (2023)
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
Hi-Light: A Path to high-fidelity, high-resolution video relighting with a Novel Evaluation Paradigm
by: Liu, Xiangrui, et al.
Published: (2026)
by: Liu, Xiangrui, et al.
Published: (2026)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
by: Ou, Siqu, et al.
Published: (2026)
by: Ou, Siqu, et al.
Published: (2026)
Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology
by: Han, Minghao, et al.
Published: (2026)
by: Han, Minghao, et al.
Published: (2026)
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering
by: Shang, Xinyi, et al.
Published: (2026)
by: Shang, Xinyi, et al.
Published: (2026)
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
by: Li, Pengteng, et al.
Published: (2025)
by: Li, Pengteng, et al.
Published: (2025)
Similar Items
-
Beyond Flatlands: Unlocking Spatial Intelligence by Decoupling 3D Reasoning from Numerical Regression
by: Guo, Zhongbin, et al.
Published: (2025) -
LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency
by: Guo, Zhongbin, et al.
Published: (2025) -
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models
by: Guo, Zhongbin, et al.
Published: (2025) -
How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
by: Yang, Zhen, et al.
Published: (2026) -
The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
by: Deng, Yonghong, et al.
Published: (2026)