CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Kaizhen, Feng, Yang, Du, Heqing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Does Visual Token Pruning Improve Calibration? An Empirical Study on Confidence in MLLMs
by: Tan, Kaizhen
Published: (2026)
by: Tan, Kaizhen
Published: (2026)
UrbanVGGT: Scalable Sidewalk Width Estimation from Street View Images
by: Tan, Kaizhen, et al.
Published: (2026)
by: Tan, Kaizhen, et al.
Published: (2026)
Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
by: Tan, Kaizhen
Published: (2025)
by: Tan, Kaizhen
Published: (2025)
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
by: Zhang, Yiming, et al.
Published: (2026)
by: Zhang, Yiming, et al.
Published: (2026)
Decoding Tourist Perception in Historic Urban Quarters with Multimodal Social Media Data: An AI-Based Framework and Evidence from Shanghai
by: Tan, Kaizhen, et al.
Published: (2025)
by: Tan, Kaizhen, et al.
Published: (2025)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
STUPD: A Synthetic Dataset for Spatial and Temporal Relation Reasoning
by: Agrawal, Palaash, et al.
Published: (2023)
by: Agrawal, Palaash, et al.
Published: (2023)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026)
by: Ma, Xueqi, et al.
Published: (2026)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
by: Zhang, Wanyue, et al.
Published: (2026)
by: Zhang, Wanyue, et al.
Published: (2026)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
Dual Attribute-Spatial Relation Alignment for 3D Visual Grounding
by: Xu, Yue, et al.
Published: (2024)
by: Xu, Yue, et al.
Published: (2024)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
ArtVLM: Attribute Recognition Through Vision-Based Prefix Language Modeling
by: Zhu, William Yicheng, et al.
Published: (2024)
by: Zhu, William Yicheng, et al.
Published: (2024)
Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation
by: He, Minggui, et al.
Published: (2026)
by: He, Minggui, et al.
Published: (2026)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
Tri-Bench: Stress-Testing VLM Reliability on Spatial Reasoning under Camera Tilt and Object Interference
by: Bendkhale, Amit
Published: (2025)
by: Bendkhale, Amit
Published: (2025)
TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
by: Ly, Vinh-Thuan, et al.
Published: (2025)
by: Ly, Vinh-Thuan, et al.
Published: (2025)
RARL: Improving Medical VLM Reasoning and Generalization with Reinforcement Learning and LoRA under Data and Hardware Constraints
by: Pham, Tan-Hanh, et al.
Published: (2025)
by: Pham, Tan-Hanh, et al.
Published: (2025)
Structured Relational Reasoning for Group Activity Assessment
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025)
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
by: Luo, Meng, et al.
Published: (2026)
by: Luo, Meng, et al.
Published: (2026)
GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
by: Liu, Jiajin, et al.
Published: (2026)
by: Liu, Jiajin, et al.
Published: (2026)
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
by: Li, Peizheng, et al.
Published: (2025)
by: Li, Peizheng, et al.
Published: (2025)
N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
SSR: Pushing the Limit of Spatial Intelligence with Structured Scene Reasoning
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
SpatialMosaic: A Multiview VLM Dataset for Partial Visibility
by: Lee, Kanghee, et al.
Published: (2025)
by: Lee, Kanghee, et al.
Published: (2025)
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
by: Bai, Xuehai, et al.
Published: (2026)
by: Bai, Xuehai, et al.
Published: (2026)
Evaluating the Generation of Spatial Relations in Text and Image Generative Models
by: Sim, Shang Hong, et al.
Published: (2024)
by: Sim, Shang Hong, et al.
Published: (2024)
VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation
by: O'Mahony, Felix, et al.
Published: (2025)
by: O'Mahony, Felix, et al.
Published: (2025)
Question-Aware Evidence Ledgers for Video Relational Reasoning
by: Ou, Yilin, et al.
Published: (2026)
by: Ou, Yilin, et al.
Published: (2026)
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
by: Huang, Zhipeng, et al.
Published: (2024)
by: Huang, Zhipeng, et al.
Published: (2024)
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
by: Xue, Xizhe, et al.
Published: (2024)
by: Xue, Xizhe, et al.
Published: (2024)
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
by: Lin, Jiawen, et al.
Published: (2025)
by: Lin, Jiawen, et al.
Published: (2025)
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
by: Qiu, Jason, et al.
Published: (2026)
by: Qiu, Jason, et al.
Published: (2026)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
by: Fateh, Fawad Javed, et al.
Published: (2024)
by: Fateh, Fawad Javed, et al.
Published: (2024)
Vision-Language Memory for Spatial Reasoning
by: Liu, Zuntao, et al.
Published: (2025)
by: Liu, Zuntao, et al.
Published: (2025)
AdaFV: Rethinking of Visual-Language alignment for VLM acceleration
by: Han, Jiayi, et al.
Published: (2025)
by: Han, Jiayi, et al.
Published: (2025)
Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge
by: Sun, Yuyang, et al.
Published: (2026)
by: Sun, Yuyang, et al.
Published: (2026)
Structured Spatial Reasoning with Open Vocabulary Object Detectors
by: Nejatishahidin, Negar, et al.
Published: (2024)
by: Nejatishahidin, Negar, et al.
Published: (2024)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
by: Wu, Zhenkai, et al.
Published: (2025)
by: Wu, Zhenkai, et al.
Published: (2025)
VERDI: VLM-Embedded Reasoning for Autonomous Driving
by: Feng, Bowen, et al.
Published: (2025)
by: Feng, Bowen, et al.
Published: (2025)
Similar Items
-
Does Visual Token Pruning Improve Calibration? An Empirical Study on Confidence in MLLMs
by: Tan, Kaizhen
Published: (2026) -
UrbanVGGT: Scalable Sidewalk Width Estimation from Street View Images
by: Tan, Kaizhen, et al.
Published: (2026) -
Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
by: Tan, Kaizhen
Published: (2025) -
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
by: Zhang, Yiming, et al.
Published: (2026) -
Decoding Tourist Perception in Historic Urban Quarters with Multimodal Social Media Data: An AI-Based Framework and Evidence from Shanghai
by: Tan, Kaizhen, et al.
Published: (2025)