3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhan, Shaoxiong, Lai, Yanlin, Liu, Zheng, Lin, Hai, Li, Shen, Cai, Xiaodong, Lin, Zijian, Huang, Wen, Zheng, Hai-Tao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Auxiliary-Hyperparameter-Free Sampling: Entropy Equilibrium for Text Generation
von: Cai, Xiaodong, et al.
Veröffentlicht: (2025)
von: Cai, Xiaodong, et al.
Veröffentlicht: (2025)
QAEA-DR: A Unified Text Augmentation Framework for Dense Retrieval
von: Tan, Hongming, et al.
Veröffentlicht: (2024)
von: Tan, Hongming, et al.
Veröffentlicht: (2024)
MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2025)
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
von: Li, Chengzu, et al.
Veröffentlicht: (2024)
von: Li, Chengzu, et al.
Veröffentlicht: (2024)
A Hierarchical Framework for Measuring Scientific Paper Innovation via Large Language Models
von: Tan, Hongming, et al.
Veröffentlicht: (2025)
von: Tan, Hongming, et al.
Veröffentlicht: (2025)
GaussianCAD: Robust Self-Supervised CAD Reconstruction from Three Orthographic Views Using 3D Gaussian Splatting
von: Zhou, Zheng, et al.
Veröffentlicht: (2025)
von: Zhou, Zheng, et al.
Veröffentlicht: (2025)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
von: Liu, Daixian, et al.
Veröffentlicht: (2026)
von: Liu, Daixian, et al.
Veröffentlicht: (2026)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
ParGo: Bridging Vision-Language with Partial and Global Views
von: Wang, An-Lan, et al.
Veröffentlicht: (2024)
von: Wang, An-Lan, et al.
Veröffentlicht: (2024)
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
von: Li, Dingming, et al.
Veröffentlicht: (2025)
von: Li, Dingming, et al.
Veröffentlicht: (2025)
Beyond Semantics: An Evidential Reasoning-Aware Multi-View Learning Framework for Trustworthy Mental Health Prediction
von: Ruan, Yucheng, et al.
Veröffentlicht: (2026)
von: Ruan, Yucheng, et al.
Veröffentlicht: (2026)
CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis
von: Li, Weijia, et al.
Veröffentlicht: (2024)
von: Li, Weijia, et al.
Veröffentlicht: (2024)
Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning
von: Li, Rui, et al.
Veröffentlicht: (2024)
von: Li, Rui, et al.
Veröffentlicht: (2024)
MindCube: Spatial Mental Modeling from Limited Views
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
Learning Multi-View Spatial Reasoning from Cross-View Relations
von: Jeong, Suchae, et al.
Veröffentlicht: (2026)
von: Jeong, Suchae, et al.
Veröffentlicht: (2026)
Tri-Perspective View Decomposition for Geometry-Aware Depth Completion
von: Yan, Zhiqiang, et al.
Veröffentlicht: (2024)
von: Yan, Zhiqiang, et al.
Veröffentlicht: (2024)
Behavioral Bias of Vision-Language Models: A Behavioral Finance View
von: Xiao, Yuhang, et al.
Veröffentlicht: (2024)
von: Xiao, Yuhang, et al.
Veröffentlicht: (2024)
Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation
von: Zhang, Hai, et al.
Veröffentlicht: (2026)
von: Zhang, Hai, et al.
Veröffentlicht: (2026)
MM-Point: Multi-View Information-Enhanced Multi-Modal Self-Supervised 3D Point Cloud Understanding
von: Yu, Hai-Tao, et al.
Veröffentlicht: (2024)
von: Yu, Hai-Tao, et al.
Veröffentlicht: (2024)
Orthographer: Generating Orthographic‐Style Projections for Elongated Architectural Structures
von: Yu‐Hsuan Hsieh, et al.
Veröffentlicht: (2026)
von: Yu‐Hsuan Hsieh, et al.
Veröffentlicht: (2026)
View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
von: Liu, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Liu, Yuanyuan, et al.
Veröffentlicht: (2025)
OpenFMNav: Towards Open-Set Zero-Shot Object Navigation via Vision-Language Foundation Models
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2024)
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
von: Fan, Qingyu, et al.
Veröffentlicht: (2026)
von: Fan, Qingyu, et al.
Veröffentlicht: (2026)
Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence
von: Zhang, Yanbing, et al.
Veröffentlicht: (2026)
von: Zhang, Yanbing, et al.
Veröffentlicht: (2026)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
von: Song, Zijian, et al.
Veröffentlicht: (2025)
von: Song, Zijian, et al.
Veröffentlicht: (2025)
Towards Cross-View Point Correspondence in Vision-Language Models
von: Wang, Yipu, et al.
Veröffentlicht: (2025)
von: Wang, Yipu, et al.
Veröffentlicht: (2025)
UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features
von: Cui, Haowang, et al.
Veröffentlicht: (2025)
von: Cui, Haowang, et al.
Veröffentlicht: (2025)
Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2024)
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2024)
Perception Compressor: A Training-Free Prompt Compression Framework in Long Context Scenarios
von: Tang, Jiwei, et al.
Veröffentlicht: (2024)
von: Tang, Jiwei, et al.
Veröffentlicht: (2024)
CoV: Chain-of-View Prompting for Spatial Reasoning
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
NeRF-Det++: Incorporating Semantic Cues and Perspective-aware Depth Supervision for Indoor Multi-View 3D Detection
von: Huang, Chenxi, et al.
Veröffentlicht: (2024)
von: Huang, Chenxi, et al.
Veröffentlicht: (2024)
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
von: Lee, Phillip Y., et al.
Veröffentlicht: (2025)
von: Lee, Phillip Y., et al.
Veröffentlicht: (2025)
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
von: Li, Hongxing, et al.
Veröffentlicht: (2025)
von: Li, Hongxing, et al.
Veröffentlicht: (2025)
View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View Selection
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images
von: Ro, Juneyoung, et al.
Veröffentlicht: (2025)
von: Ro, Juneyoung, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Auxiliary-Hyperparameter-Free Sampling: Entropy Equilibrium for Text Generation
von: Cai, Xiaodong, et al.
Veröffentlicht: (2025) -
QAEA-DR: A Unified Text Augmentation Framework for Dense Retrieval
von: Tan, Hongming, et al.
Veröffentlicht: (2024) -
MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2025) -
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
von: Tao, Xingjian, et al.
Veröffentlicht: (2026) -
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
von: Li, Chengzu, et al.
Veröffentlicht: (2024)