SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Batra, Hunar, Tu, Haoqin, Chen, Hardy, Lin, Yuanze, Xie, Cihang, Clark, Ronald |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ViLBench: A Suite for Vision-Language Process Reward Modeling
von: Tu, Haoqin, et al.
Veröffentlicht: (2025)
von: Tu, Haoqin, et al.
Veröffentlicht: (2025)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
Towards Understanding Multimodal Fine-Tuning: Spatial Features
von: Naghashyar, Lachin, et al.
Veröffentlicht: (2026)
von: Naghashyar, Lachin, et al.
Veröffentlicht: (2026)
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
von: Daxberger, Erik, et al.
Veröffentlicht: (2025)
von: Daxberger, Erik, et al.
Veröffentlicht: (2025)
MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following
von: Lee, Jaeyun, et al.
Veröffentlicht: (2026)
von: Lee, Jaeyun, et al.
Veröffentlicht: (2026)
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
von: Chen, Hardy, et al.
Veröffentlicht: (2025)
von: Chen, Hardy, et al.
Veröffentlicht: (2025)
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation
von: Wang, Zijun, et al.
Veröffentlicht: (2024)
von: Wang, Zijun, et al.
Veröffentlicht: (2024)
Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
von: Wu, Juncheng, et al.
Veröffentlicht: (2025)
von: Wu, Juncheng, et al.
Veröffentlicht: (2025)
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
von: Rajabi, Navid, et al.
Veröffentlicht: (2024)
von: Rajabi, Navid, et al.
Veröffentlicht: (2024)
DreamPolisher: Towards High-Quality Text-to-3D Generation via Geometric Diffusion
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
von: He, Jixuan, et al.
Veröffentlicht: (2026)
von: He, Jixuan, et al.
Veröffentlicht: (2026)
Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
von: Ma, Chuang, et al.
Veröffentlicht: (2026)
von: Ma, Chuang, et al.
Veröffentlicht: (2026)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
What If We Recaption Billions of Web Images with LLaMA-3?
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
von: Ding, Shengyuan, et al.
Veröffentlicht: (2025)
von: Ding, Shengyuan, et al.
Veröffentlicht: (2025)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
von: Li, Zhuowan, et al.
Veröffentlicht: (2022)
von: Li, Zhuowan, et al.
Veröffentlicht: (2022)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
Olympus: A Universal Task Router for Computer Vision Tasks
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
von: Liu, Tianhui, et al.
Veröffentlicht: (2026)
von: Liu, Tianhui, et al.
Veröffentlicht: (2026)
MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
von: Liu, Daixian, et al.
Veröffentlicht: (2026)
von: Liu, Daixian, et al.
Veröffentlicht: (2026)
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
von: Taguchi, Shun, et al.
Veröffentlicht: (2025)
von: Taguchi, Shun, et al.
Veröffentlicht: (2025)
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2026)
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2026)
Controlling Multimodal LLMs via Reward-guided Decoding
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments
von: Li, Dinging, et al.
Veröffentlicht: (2026)
von: Li, Dinging, et al.
Veröffentlicht: (2026)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
von: Li, Hongxing, et al.
Veröffentlicht: (2025)
von: Li, Hongxing, et al.
Veröffentlicht: (2025)
ProxyThinker: Test-Time Guidance through Small Visual Reasoners
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
von: Guo, Jiajie, et al.
Veröffentlicht: (2025)
von: Guo, Jiajie, et al.
Veröffentlicht: (2025)
11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ViLBench: A Suite for Vision-Language Process Reward Modeling
von: Tu, Haoqin, et al.
Veröffentlicht: (2025) -
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026) -
Towards Understanding Multimodal Fine-Tuning: Spatial Features
von: Naghashyar, Lachin, et al.
Veröffentlicht: (2026) -
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
von: Wu, Juncheng, et al.
Veröffentlicht: (2026) -
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
von: Daxberger, Erik, et al.
Veröffentlicht: (2025)