Spatial Preference Rewarding for MLLMs Spatial Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Han, Gao, Peng, Lu, Lewei, Zhang, Xiaoqin, Shao, Ling, Lu, Shijian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Masked AutoDecoder is Effective Multi-Task Vision Generalist
by: Qiu, Han, et al.
Published: (2024)
by: Qiu, Han, et al.
Published: (2024)
Multimodal 3D Reasoning Segmentation with Complex Scenes
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
by: Qiu, Han, et al.
Published: (2024)
by: Qiu, Han, et al.
Published: (2024)
Weakly Supervised Monocular 3D Detection with a Single-View Image
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
A Survey of Label-Efficient Deep Learning for 3D Point Clouds
by: Xiao, Aoran, et al.
Published: (2023)
by: Xiao, Aoran, et al.
Published: (2023)
L3DR: 3D-aware LiDAR Diffusion and Rectification
by: Liu, Quan, et al.
Published: (2026)
by: Liu, Quan, et al.
Published: (2026)
On the Generalization Capacities of MLLMs for Spatial Intelligence
by: Zhang, Gongjie, et al.
Published: (2026)
by: Zhang, Gongjie, et al.
Published: (2026)
STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-Regularizations
by: Wei, Yu, et al.
Published: (2025)
by: Wei, Yu, et al.
Published: (2025)
ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance
by: Li, Duo, et al.
Published: (2025)
by: Li, Duo, et al.
Published: (2025)
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
by: Li, Duo, et al.
Published: (2025)
by: Li, Duo, et al.
Published: (2025)
MuSASplat: Efficient Sparse-View 3D Gaussian Splats via Lightweight Multi-Scale Adaptation
by: Xu, Muyu, et al.
Published: (2025)
by: Xu, Muyu, et al.
Published: (2025)
MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked Autoencoders
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
Exploring 3D Reasoning-Driven Planning: From Implicit Human Intentions to Route-Aware Activity Planning
by: Jiang, Xueying, et al.
Published: (2025)
by: Jiang, Xueying, et al.
Published: (2025)
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
by: Chang, Chun-Peng, et al.
Published: (2024)
by: Chang, Chun-Peng, et al.
Published: (2024)
SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
by: Wang, Peiyao, et al.
Published: (2025)
by: Wang, Peiyao, et al.
Published: (2025)
Learning to Prompt Segment Anything Models
by: Huang, Jiaxing, et al.
Published: (2024)
by: Huang, Jiaxing, et al.
Published: (2024)
Historical Test-time Prompt Tuning for Vision Foundation Models
by: Zhang, Jingyi, et al.
Published: (2024)
by: Zhang, Jingyi, et al.
Published: (2024)
Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
Novel View Extrapolation with Video Diffusion Priors
by: Liu, Kunhao, et al.
Published: (2024)
by: Liu, Kunhao, et al.
Published: (2024)
LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors
by: Jin, Sheng, et al.
Published: (2024)
by: Jin, Sheng, et al.
Published: (2024)
Visual Jigsaw Post-Training Improves MLLMs
by: Wu, Penghao, et al.
Published: (2025)
by: Wu, Penghao, et al.
Published: (2025)
SpatialTree: How Spatial Abilities Branch Out in MLLMs
by: Xiao, Yuxi, et al.
Published: (2025)
by: Xiao, Yuxi, et al.
Published: (2025)
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
SOGS: Second-Order Anchor for Advanced 3D Gaussian Splatting
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
CAT-SAM: Conditional Tuning for Few-Shot Adaptation of Segment Anything Model
by: Xiao, Aoran, et al.
Published: (2024)
by: Xiao, Aoran, et al.
Published: (2024)
Data-Efficient Generalization for Zero-shot Composed Image Retrieval
by: Chen, Zining, et al.
Published: (2025)
by: Chen, Zining, et al.
Published: (2025)
PacGDC: Label-Efficient Generalizable Depth Completion with Projection Ambiguity and Consistency
by: Wang, Haotian, et al.
Published: (2025)
by: Wang, Haotian, et al.
Published: (2025)
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
by: Zhang, Wanyue, et al.
Published: (2025)
by: Zhang, Wanyue, et al.
Published: (2025)
Modeling Continuous Motion for 3D Point Cloud Object Tracking
by: Luo, Zhipeng, et al.
Published: (2023)
by: Luo, Zhipeng, et al.
Published: (2023)
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024)
by: Zhang, Mengxi, et al.
Published: (2024)
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
Enhancing Spatial Understanding in Image Generation via Reward Modeling
by: Tang, Zhenyu, et al.
Published: (2026)
by: Tang, Zhenyu, et al.
Published: (2026)
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
by: Wang, Siting, et al.
Published: (2025)
by: Wang, Siting, et al.
Published: (2025)
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
by: Jiang, Xueying, et al.
Published: (2026)
by: Jiang, Xueying, et al.
Published: (2026)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning
by: Long, Yancheng, et al.
Published: (2026)
by: Long, Yancheng, et al.
Published: (2026)
Spatial-ORMLLM: Improve Spatial Relation Understanding in the Operating Room with Multimodal Large Language Model
by: He, Peiqi, et al.
Published: (2025)
by: He, Peiqi, et al.
Published: (2025)
Similar Items
-
Masked AutoDecoder is Effective Multi-Task Vision Generalist
by: Qiu, Han, et al.
Published: (2024) -
Multimodal 3D Reasoning Segmentation with Complex Scenes
by: Jiang, Xueying, et al.
Published: (2024) -
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
by: Qiu, Han, et al.
Published: (2024) -
Weakly Supervised Monocular 3D Detection with a Single-View Image
by: Jiang, Xueying, et al.
Published: (2024) -
A Survey of Label-Efficient Deep Learning for 3D Point Clouds
by: Xiao, Aoran, et al.
Published: (2023)