SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Ouyang, Kun, Liu, Yuanxin, Wu, Haoning, Liu, Yi, Zhou, Hao, Zhou, Jie, Meng, Fandong, Sun, Xu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
by: Ouyang, Kun, et al.
Published: (2024)
by: Ouyang, Kun, et al.
Published: (2024)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
Video-R1: Reinforcing Video Reasoning in MLLMs
by: Feng, Kaituo, et al.
Published: (2025)
by: Feng, Kaituo, et al.
Published: (2025)
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
by: Xu, Zhiyu, et al.
Published: (2026)
by: Xu, Zhiyu, et al.
Published: (2026)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
Touch-R1: Reinforcing Touch Reasoning in MLLMs
by: Lai, Yingxin, et al.
Published: (2026)
by: Lai, Yingxin, et al.
Published: (2026)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
by: Fan, Kaixuan, et al.
Published: (2025)
by: Fan, Kaixuan, et al.
Published: (2025)
Continuous Visual Autoregressive Generation via Score Maximization
by: Shao, Chenze, et al.
Published: (2025)
by: Shao, Chenze, et al.
Published: (2025)
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
by: Meng, Desen, et al.
Published: (2025)
by: Meng, Desen, et al.
Published: (2025)
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
by: Yu, Fangxu, et al.
Published: (2026)
by: Yu, Fangxu, et al.
Published: (2026)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
by: Wang, Peiyao, et al.
Published: (2025)
by: Wang, Peiyao, et al.
Published: (2025)
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
by: Liu, Juntao, et al.
Published: (2025)
by: Liu, Juntao, et al.
Published: (2025)
Generative Frame Sampler for Long Video Understanding
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
Explore the Hallucination on Low-level Perception for MLLMs
by: Sun, Yinan, et al.
Published: (2024)
by: Sun, Yinan, et al.
Published: (2024)
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
by: Wang, Haicheng, et al.
Published: (2026)
by: Wang, Haicheng, et al.
Published: (2026)
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models
by: Wei, Yuancheng, et al.
Published: (2026)
by: Wei, Yuancheng, et al.
Published: (2026)
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
by: Zhang, Bob, et al.
Published: (2025)
by: Zhang, Bob, et al.
Published: (2025)
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
by: Gu, Bo, et al.
Published: (2026)
by: Gu, Bo, et al.
Published: (2026)
SpatialTree: How Spatial Abilities Branch Out in MLLMs
by: Xiao, Yuxi, et al.
Published: (2025)
by: Xiao, Yuxi, et al.
Published: (2025)
Exploring the Design Space of Visual Context Representation in Video MLLMs
by: Du, Yifan, et al.
Published: (2024)
by: Du, Yifan, et al.
Published: (2024)
Wan-R1: Verifiable-Reinforcement Learning for Video Reasoning
by: Liu, Ming, et al.
Published: (2026)
by: Liu, Ming, et al.
Published: (2026)
Temporal Reasoning Transfer from Text to Video
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
Low-Light Video Enhancement via Spatial-Temporal Consistent Decomposition
by: Xu, Xiaogang, et al.
Published: (2024)
by: Xu, Xiaogang, et al.
Published: (2024)
Low-Light Video Enhancement with An Effective Spatial-Temporal Decomposition Paradigm
by: Xu, Xiaogang, et al.
Published: (2026)
by: Xu, Xiaogang, et al.
Published: (2026)
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
ViSS-R1: Self-Supervised Reinforcement Video Reasoning
by: Fang, Bo, et al.
Published: (2025)
by: Fang, Bo, et al.
Published: (2025)
TempCompass: Do Video LLMs Really Understand Videos?
by: Liu, Yuanxin, et al.
Published: (2024)
by: Liu, Yuanxin, et al.
Published: (2024)
MaskMamba: A Hybrid Mamba-Transformer Model for Masked Image Generation
by: Chen, Wenchao, et al.
Published: (2024)
by: Chen, Wenchao, et al.
Published: (2024)
D2C: Unlocking the Potential of Continuous Autoregressive Image Generation with Discrete Tokens
by: Wang, Panpan, et al.
Published: (2025)
by: Wang, Panpan, et al.
Published: (2025)
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
Reinforcing Consistency in Video MLLMs with Structured Rewards
by: Quan, Yihao, et al.
Published: (2026)
by: Quan, Yihao, et al.
Published: (2026)
Spatial-Conditioned Reasoning in Long-Egocentric Videos
by: Tribble, James, et al.
Published: (2026)
by: Tribble, James, et al.
Published: (2026)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
by: Sun, Peiwen, et al.
Published: (2025)
by: Sun, Peiwen, et al.
Published: (2025)
When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?
by: Ye, Qilang, et al.
Published: (2025)
by: Ye, Qilang, et al.
Published: (2025)
TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
by: Li, Shicheng, et al.
Published: (2025)
by: Li, Shicheng, et al.
Published: (2025)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
Similar Items
-
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
by: Ouyang, Kun, et al.
Published: (2024) -
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
by: Liu, Yuanxin, et al.
Published: (2025) -
Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
by: Ouyang, Kun, et al.
Published: (2025) -
Video-R1: Reinforcing Video Reasoning in MLLMs
by: Feng, Kaituo, et al.
Published: (2025) -
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
by: Xu, Zhiyu, et al.
Published: (2026)