3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhan, Shaoxiong, Lai, Yanlin, Liu, Zheng, Lin, Hai, Li, Shen, Cai, Xiaodong, Lin, Zijian, Huang, Wen, Zheng, Hai-Tao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Auxiliary-Hyperparameter-Free Sampling: Entropy Equilibrium for Text Generation
by: Cai, Xiaodong, et al.
Published: (2025)
by: Cai, Xiaodong, et al.
Published: (2025)
QAEA-DR: A Unified Text Augmentation Framework for Dense Retrieval
by: Tan, Hongming, et al.
Published: (2024)
by: Tan, Hongming, et al.
Published: (2024)
MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
by: Zhan, Shaoxiong, et al.
Published: (2025)
by: Zhan, Shaoxiong, et al.
Published: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
by: Tao, Xingjian, et al.
Published: (2026)
by: Tao, Xingjian, et al.
Published: (2026)
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
by: Li, Chengzu, et al.
Published: (2024)
by: Li, Chengzu, et al.
Published: (2024)
A Hierarchical Framework for Measuring Scientific Paper Innovation via Large Language Models
by: Tan, Hongming, et al.
Published: (2025)
by: Tan, Hongming, et al.
Published: (2025)
GaussianCAD: Robust Self-Supervised CAD Reconstruction from Three Orthographic Views Using 3D Gaussian Splatting
by: Zhou, Zheng, et al.
Published: (2025)
by: Zhou, Zheng, et al.
Published: (2025)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
by: Wu, Shunlong, et al.
Published: (2026)
by: Wu, Shunlong, et al.
Published: (2026)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
by: Liu, Fanfan, et al.
Published: (2024)
by: Liu, Fanfan, et al.
Published: (2024)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025)
by: Feng, Zhiyuan, et al.
Published: (2025)
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
by: Liu, Daixian, et al.
Published: (2026)
by: Liu, Daixian, et al.
Published: (2026)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
by: Gholami, Mohsen, et al.
Published: (2025)
by: Gholami, Mohsen, et al.
Published: (2025)
ParGo: Bridging Vision-Language with Partial and Global Views
by: Wang, An-Lan, et al.
Published: (2024)
by: Wang, An-Lan, et al.
Published: (2024)
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
by: Li, Dingming, et al.
Published: (2025)
by: Li, Dingming, et al.
Published: (2025)
Beyond Semantics: An Evidential Reasoning-Aware Multi-View Learning Framework for Trustworthy Mental Health Prediction
by: Ruan, Yucheng, et al.
Published: (2026)
by: Ruan, Yucheng, et al.
Published: (2026)
CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis
by: Li, Weijia, et al.
Published: (2024)
by: Li, Weijia, et al.
Published: (2024)
Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
MindCube: Spatial Mental Modeling from Limited Views
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
by: Tang, Yihong, et al.
Published: (2024)
by: Tang, Yihong, et al.
Published: (2024)
Learning Multi-View Spatial Reasoning from Cross-View Relations
by: Jeong, Suchae, et al.
Published: (2026)
by: Jeong, Suchae, et al.
Published: (2026)
Tri-Perspective View Decomposition for Geometry-Aware Depth Completion
by: Yan, Zhiqiang, et al.
Published: (2024)
by: Yan, Zhiqiang, et al.
Published: (2024)
Behavioral Bias of Vision-Language Models: A Behavioral Finance View
by: Xiao, Yuhang, et al.
Published: (2024)
by: Xiao, Yuhang, et al.
Published: (2024)
Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation
by: Zhang, Hai, et al.
Published: (2026)
by: Zhang, Hai, et al.
Published: (2026)
MM-Point: Multi-View Information-Enhanced Multi-Modal Self-Supervised 3D Point Cloud Understanding
by: Yu, Hai-Tao, et al.
Published: (2024)
by: Yu, Hai-Tao, et al.
Published: (2024)
Orthographer: Generating Orthographic‐Style Projections for Elongated Architectural Structures
by: Yu‐Hsuan Hsieh, et al.
Published: (2026)
by: Yu‐Hsuan Hsieh, et al.
Published: (2026)
View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
by: Liu, Yuanyuan, et al.
Published: (2025)
by: Liu, Yuanyuan, et al.
Published: (2025)
OpenFMNav: Towards Open-Set Zero-Shot Object Navigation via Vision-Language Foundation Models
by: Kuang, Yuxuan, et al.
Published: (2024)
by: Kuang, Yuxuan, et al.
Published: (2024)
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
by: Fan, Qingyu, et al.
Published: (2026)
by: Fan, Qingyu, et al.
Published: (2026)
Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence
by: Zhang, Yanbing, et al.
Published: (2026)
by: Zhang, Yanbing, et al.
Published: (2026)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
Towards Cross-View Point Correspondence in Vision-Language Models
by: Wang, Yipu, et al.
Published: (2025)
by: Wang, Yipu, et al.
Published: (2025)
UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features
by: Cui, Haowang, et al.
Published: (2025)
by: Cui, Haowang, et al.
Published: (2025)
Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin
by: Lin, Pin-Jie, et al.
Published: (2024)
by: Lin, Pin-Jie, et al.
Published: (2024)
Perception Compressor: A Training-Free Prompt Compression Framework in Long Context Scenarios
by: Tang, Jiwei, et al.
Published: (2024)
by: Tang, Jiwei, et al.
Published: (2024)
CoV: Chain-of-View Prompting for Spatial Reasoning
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
NeRF-Det++: Incorporating Semantic Cues and Perspective-aware Depth Supervision for Indoor Multi-View 3D Detection
by: Huang, Chenxi, et al.
Published: (2024)
by: Huang, Chenxi, et al.
Published: (2024)
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
by: Lee, Phillip Y., et al.
Published: (2025)
by: Lee, Phillip Y., et al.
Published: (2025)
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
by: Li, Hongxing, et al.
Published: (2025)
by: Li, Hongxing, et al.
Published: (2025)
View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View Selection
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images
by: Ro, Juneyoung, et al.
Published: (2025)
by: Ro, Juneyoung, et al.
Published: (2025)
Similar Items
-
Auxiliary-Hyperparameter-Free Sampling: Entropy Equilibrium for Text Generation
by: Cai, Xiaodong, et al.
Published: (2025) -
QAEA-DR: A Unified Text Augmentation Framework for Dense Retrieval
by: Tan, Hongming, et al.
Published: (2024) -
MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
by: Zhan, Shaoxiong, et al.
Published: (2025) -
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
by: Tao, Xingjian, et al.
Published: (2026) -
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
by: Li, Chengzu, et al.
Published: (2024)