GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zhaochen, Qiao, Limeng, Wan, Guanglu, Jiang, Tingting |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PLUG: Revisiting Amodal Segmentation with Foundation Model and Hierarchical Focus
by: Liu, Zhaochen, et al.
Published: (2024)
by: Liu, Zhaochen, et al.
Published: (2024)
Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
by: Liu, Zhaochen, et al.
Published: (2025)
by: Liu, Zhaochen, et al.
Published: (2025)
UniComp: Rethinking Video Compression Through Informational Uniqueness
by: Yuan, Chao, et al.
Published: (2025)
by: Yuan, Chao, et al.
Published: (2025)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
by: Xu, Runsen, et al.
Published: (2025)
by: Xu, Runsen, et al.
Published: (2025)
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
by: Liu, Runzhou, et al.
Published: (2026)
by: Liu, Runzhou, et al.
Published: (2026)
BLADE: Box-Level Supervised Amodal Segmentation through Directed Expansion
by: Liu, Zhaochen, et al.
Published: (2024)
by: Liu, Zhaochen, et al.
Published: (2024)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations
by: Yuan, Jiangye, et al.
Published: (2026)
by: Yuan, Jiangye, et al.
Published: (2026)
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
Judge Anything: MLLM as a Judge Across Any Modality
by: Pu, Shu, et al.
Published: (2025)
by: Pu, Shu, et al.
Published: (2025)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
by: Li, YuQian, et al.
Published: (2025)
by: Li, YuQian, et al.
Published: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
X2SAM: Any Segmentation in Images and Videos
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments
by: Li, Dinging, et al.
Published: (2026)
by: Li, Dinging, et al.
Published: (2026)
Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts
by: Zhu, Zhihao, et al.
Published: (2026)
by: Zhu, Zhihao, et al.
Published: (2026)
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
by: Xia, Yinan, et al.
Published: (2025)
by: Xia, Yinan, et al.
Published: (2025)
Amodal Segmentation for Laparoscopic Surgery Video Instruments
by: Shi, Ruohua, et al.
Published: (2024)
by: Shi, Ruohua, et al.
Published: (2024)
Effective Training Data Synthesis for Improving MLLM Chart Understanding
by: Yang, Yuwei, et al.
Published: (2025)
by: Yang, Yuwei, et al.
Published: (2025)
SGAD: Semantic and Geometric-aware Descriptor for Local Feature Matching
by: Liu, Xiangzeng, et al.
Published: (2025)
by: Liu, Xiangzeng, et al.
Published: (2025)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
by: Li, Yuying, et al.
Published: (2025)
by: Li, Yuying, et al.
Published: (2025)
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
by: Gao, Timin, et al.
Published: (2024)
by: Gao, Timin, et al.
Published: (2024)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder
by: Cho, Seunghyuk, et al.
Published: (2025)
by: Cho, Seunghyuk, et al.
Published: (2025)
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
by: Huang, Han, et al.
Published: (2024)
by: Huang, Han, et al.
Published: (2024)
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
by: Ma, Chuang, et al.
Published: (2026)
by: Ma, Chuang, et al.
Published: (2026)
Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
by: Pi, Renjie, et al.
Published: (2024)
by: Pi, Renjie, et al.
Published: (2024)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
by: Hua, Jiacheng, et al.
Published: (2026)
by: Hua, Jiacheng, et al.
Published: (2026)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
by: Xu, Binqian, et al.
Published: (2024)
by: Xu, Binqian, et al.
Published: (2024)
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
by: Slyman, Eric, et al.
Published: (2025)
by: Slyman, Eric, et al.
Published: (2025)
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
by: Cai, Shihao, et al.
Published: (2024)
by: Cai, Shihao, et al.
Published: (2024)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
by: Liang, Yiming, et al.
Published: (2026)
by: Liang, Yiming, et al.
Published: (2026)
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
by: Xiao, Kelaiti, et al.
Published: (2025)
by: Xiao, Kelaiti, et al.
Published: (2025)
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
by: Kuang, Jiayi, et al.
Published: (2024)
by: Kuang, Jiayi, et al.
Published: (2024)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
by: Tang, Yihong, et al.
Published: (2024)
by: Tang, Yihong, et al.
Published: (2024)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
by: Wang, Chuhan, et al.
Published: (2026)
by: Wang, Chuhan, et al.
Published: (2026)
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
by: Liao, Yuan-Hong, et al.
Published: (2024)
by: Liao, Yuan-Hong, et al.
Published: (2024)
Similar Items
-
PLUG: Revisiting Amodal Segmentation with Foundation Model and Hierarchical Focus
by: Liu, Zhaochen, et al.
Published: (2024) -
Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
by: Liu, Zhaochen, et al.
Published: (2025) -
UniComp: Rethinking Video Compression Through Informational Uniqueness
by: Yuan, Chao, et al.
Published: (2025) -
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
by: Xu, Runsen, et al.
Published: (2025) -
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
by: Liu, Runzhou, et al.
Published: (2026)