Spatial-ORMLLM: Improve Spatial Relation Understanding in the Operating Room with Multimodal Large Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Peiqi, Zhang, Zhenhao, Zhang, Yixiang, Zhao, Xiongjun, Peng, Shaoliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Multimodal Large Language Models Understand Spatial Relations?
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
SpatialBot: Precise Spatial Understanding with Vision Language Models
von: Cai, Wenxiao, et al.
Veröffentlicht: (2024)
von: Cai, Wenxiao, et al.
Veröffentlicht: (2024)
SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
von: Wang, Hanqing, et al.
Veröffentlicht: (2025)
von: Wang, Hanqing, et al.
Veröffentlicht: (2025)
Spatial Preference Rewarding for MLLMs Spatial Understanding
von: Qiu, Han, et al.
Veröffentlicht: (2025)
von: Qiu, Han, et al.
Veröffentlicht: (2025)
SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
von: Guo, Xianda, et al.
Veröffentlicht: (2024)
von: Guo, Xianda, et al.
Veröffentlicht: (2024)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
von: Dongfang, Zihao, et al.
Veröffentlicht: (2025)
von: Dongfang, Zihao, et al.
Veröffentlicht: (2025)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
von: Liang, Huizhi, et al.
Veröffentlicht: (2026)
von: Liang, Huizhi, et al.
Veröffentlicht: (2026)
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
von: Ma, Chuang, et al.
Veröffentlicht: (2026)
von: Ma, Chuang, et al.
Veröffentlicht: (2026)
Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation
von: Ning, Zhenhua, et al.
Veröffentlicht: (2025)
von: Ning, Zhenhua, et al.
Veröffentlicht: (2025)
SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
von: Xu, Runsen, et al.
Veröffentlicht: (2025)
von: Xu, Runsen, et al.
Veröffentlicht: (2025)
MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving
von: Zhang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhiyuan, et al.
Veröffentlicht: (2025)
Temporal Gains, Spatial Costs: Revisiting Video Fine-Tuning in Multimodal Large Language Models
von: Zhang, Linghao, et al.
Veröffentlicht: (2026)
von: Zhang, Linghao, et al.
Veröffentlicht: (2026)
Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
Temporal-Spatial Object Relations Modeling for Vision-and-Language Navigation
von: Huang, Bowen, et al.
Veröffentlicht: (2024)
von: Huang, Bowen, et al.
Veröffentlicht: (2024)
OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
von: Zhang, Zhenhao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenhao, et al.
Veröffentlicht: (2025)
MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
von: Deng, Nianchen, et al.
Veröffentlicht: (2025)
von: Deng, Nianchen, et al.
Veröffentlicht: (2025)
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
von: Mirjalili, Vahid, et al.
Veröffentlicht: (2025)
von: Mirjalili, Vahid, et al.
Veröffentlicht: (2025)
NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
von: Xu, Wei, et al.
Veröffentlicht: (2025)
von: Xu, Wei, et al.
Veröffentlicht: (2025)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
von: Zhang, Weichen, et al.
Veröffentlicht: (2025)
von: Zhang, Weichen, et al.
Veröffentlicht: (2025)
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
von: Chen, Wei, et al.
Veröffentlicht: (2026)
von: Chen, Wei, et al.
Veröffentlicht: (2026)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
EventGPT: Event Stream Understanding with Multimodal Large Language Models
von: Liu, Shaoyu, et al.
Veröffentlicht: (2024)
von: Liu, Shaoyu, et al.
Veröffentlicht: (2024)
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
PARSE: Part-Aware Relational Spatial Modeling
von: Bai, Yinuo, et al.
Veröffentlicht: (2026)
von: Bai, Yinuo, et al.
Veröffentlicht: (2026)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
von: Yang, Fan, et al.
Veröffentlicht: (2026)
von: Yang, Fan, et al.
Veröffentlicht: (2026)
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
von: Huang, Zhipeng, et al.
Veröffentlicht: (2024)
von: Huang, Zhipeng, et al.
Veröffentlicht: (2024)
ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming
von: Zhuang, Jiedong, et al.
Veröffentlicht: (2024)
von: Zhuang, Jiedong, et al.
Veröffentlicht: (2024)
TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies?
von: Bao, Zhongyuan, et al.
Veröffentlicht: (2025)
von: Bao, Zhongyuan, et al.
Veröffentlicht: (2025)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
von: Sinha, Sanchit, et al.
Veröffentlicht: (2026)
von: Sinha, Sanchit, et al.
Veröffentlicht: (2026)
Improving Large Vision-Language Models' Understanding for Flow Field Data
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Can Multimodal Large Language Models Understand Spatial Relations?
von: Liu, Jingping, et al.
Veröffentlicht: (2025) -
SpatialBot: Precise Spatial Understanding with Vision Language Models
von: Cai, Wenxiao, et al.
Veröffentlicht: (2024) -
SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
von: Wang, Hanqing, et al.
Veröffentlicht: (2025) -
Spatial Preference Rewarding for MLLMs Spatial Understanding
von: Qiu, Han, et al.
Veröffentlicht: (2025) -
SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
von: Guo, Xianda, et al.
Veröffentlicht: (2024)