TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Chengzu, Zhang, Caiqi, Zhou, Han, Collier, Nigel, Korhonen, Anna, Vulić, Ivan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visual Planning: Let's Think Only with Images
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
Lost in Embeddings: Information Loss in Vision-Language Models
von: Li, Wenyan, et al.
Veröffentlicht: (2025)
von: Li, Wenyan, et al.
Veröffentlicht: (2025)
Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models
von: Zhou, Ej, et al.
Veröffentlicht: (2025)
von: Zhou, Ej, et al.
Veröffentlicht: (2025)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2026)
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2026)
Large Language Models are Miscalibrated In-Context Learners
von: Li, Chengzu, et al.
Veröffentlicht: (2023)
von: Li, Chengzu, et al.
Veröffentlicht: (2023)
Translation-Enhanced Multilingual Text-to-Image Generation
von: Li, Yaoyiran, et al.
Veröffentlicht: (2023)
von: Li, Yaoyiran, et al.
Veröffentlicht: (2023)
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
von: Zhou, Han, et al.
Veröffentlicht: (2024)
von: Zhou, Han, et al.
Veröffentlicht: (2024)
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
von: Hoehing, Nils, et al.
Veröffentlicht: (2025)
von: Hoehing, Nils, et al.
Veröffentlicht: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation
von: Zhong, Linqing, et al.
Veröffentlicht: (2024)
von: Zhong, Linqing, et al.
Veröffentlicht: (2024)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
von: Li, Dingming, et al.
Veröffentlicht: (2025)
von: Li, Dingming, et al.
Veröffentlicht: (2025)
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
von: Rajabi, Navid, et al.
Veröffentlicht: (2023)
von: Rajabi, Navid, et al.
Veröffentlicht: (2023)
Top2Pano: Learning to Generate Indoor Panoramas from Top-Down View
von: Zhang, Zitong, et al.
Veröffentlicht: (2025)
von: Zhang, Zitong, et al.
Veröffentlicht: (2025)
Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
von: Liu, Yinhong, et al.
Veröffentlicht: (2024)
von: Liu, Yinhong, et al.
Veröffentlicht: (2024)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model
von: Li, Ling, et al.
Veröffentlicht: (2024)
von: Li, Ling, et al.
Veröffentlicht: (2024)
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
Sherlock: Self-Correcting Reasoning in Vision-Language Models
von: Ding, Yi, et al.
Veröffentlicht: (2025)
von: Ding, Yi, et al.
Veröffentlicht: (2025)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
Improving Word Translation via Two-Stage Contrastive Learning
von: Li, Yaoyiran, et al.
Veröffentlicht: (2022)
von: Li, Yaoyiran, et al.
Veröffentlicht: (2022)
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
von: Liao, Yuan-Hong, et al.
Veröffentlicht: (2024)
von: Liao, Yuan-Hong, et al.
Veröffentlicht: (2024)
SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?
von: Wasi, Azmine Toushik, et al.
Veröffentlicht: (2026)
von: Wasi, Azmine Toushik, et al.
Veröffentlicht: (2026)
Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models
von: Qiu, Yifu, et al.
Veröffentlicht: (2025)
von: Qiu, Yifu, et al.
Veröffentlicht: (2025)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
MindCube: Spatial Mental Modeling from Limited Views
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
Representation Engineering: A Top-Down Approach to AI Transparency
von: Zou, Andy, et al.
Veröffentlicht: (2023)
von: Zou, Andy, et al.
Veröffentlicht: (2023)
Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
von: Lian, Shijie, et al.
Veröffentlicht: (2025)
von: Lian, Shijie, et al.
Veröffentlicht: (2025)
Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
von: Li, Hongxing, et al.
Veröffentlicht: (2025)
von: Li, Hongxing, et al.
Veröffentlicht: (2025)
PROGRESSLM: Towards Progress Reasoning in Vision-Language Models
von: Zhang, Jianshu, et al.
Veröffentlicht: (2026)
von: Zhang, Jianshu, et al.
Veröffentlicht: (2026)
TopKD: Top-scaled Knowledge Distillation
von: Wang, Qi, et al.
Veröffentlicht: (2025)
von: Wang, Qi, et al.
Veröffentlicht: (2025)
Vision-Language Models Can Self-Improve Reasoning via Reflection
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2024)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2024)
Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies
von: Pathak, Surendra, et al.
Veröffentlicht: (2026)
von: Pathak, Surendra, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Visual Planning: Let's Think Only with Images
von: Xu, Yi, et al.
Veröffentlicht: (2025) -
Lost in Embeddings: Information Loss in Vision-Language Models
von: Li, Wenyan, et al.
Veröffentlicht: (2025) -
Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models
von: Zhou, Ej, et al.
Veröffentlicht: (2025) -
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
von: Li, Chengzu, et al.
Veröffentlicht: (2026) -
11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
von: Li, Chengzu, et al.
Veröffentlicht: (2025)