Learning Proposes, Geometry Disposes: A Modular Framework for Efficient Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Haichao, Yang, Zhaorui, Zhang, Qian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GravCal: Single-Image Calibration of IMU Gravity Priors with Per-Sample Confidence
by: Zhu, Haichao, et al.
Published: (2026)
by: Zhu, Haichao, et al.
Published: (2026)
Ultron: Enabling Temporal Geometry Compression of 3D Mesh Sequences using Temporal Correspondence and Mesh Deformation
by: Zhu, Haichao
Published: (2024)
by: Zhu, Haichao
Published: (2024)
Sharpening Your Density Fields: Spiking Neuron Aided Fast Geometry Learning
by: Gu, Yi, et al.
Published: (2024)
by: Gu, Yi, et al.
Published: (2024)
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
by: Li, Haoyuan, et al.
Published: (2026)
by: Li, Haoyuan, et al.
Published: (2026)
SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
by: Guo, Jiajie, et al.
Published: (2025)
by: Guo, Jiajie, et al.
Published: (2025)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
by: Li, Yian, et al.
Published: (2026)
by: Li, Yian, et al.
Published: (2026)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
Make Geometry Matter for Spatial Reasoning
by: Zhang, Shihua, et al.
Published: (2026)
by: Zhang, Shihua, et al.
Published: (2026)
Efficient Bayer-Domain Video Computer Vision with Fast Motion Estimation and Learned Perception Residual
by: Wang, Haichao, et al.
Published: (2025)
by: Wang, Haichao, et al.
Published: (2025)
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
by: Zhu, Fangrui, et al.
Published: (2025)
by: Zhu, Fangrui, et al.
Published: (2025)
GraphMAR: Geometry-Aware Graph Learning Framework for Spatially Adaptive CT Metal Artifact Reduction
by: Li, Zilong, et al.
Published: (2026)
by: Li, Zilong, et al.
Published: (2026)
Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning
by: Shi, Jian, et al.
Published: (2026)
by: Shi, Jian, et al.
Published: (2026)
RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
by: Cai, Zhongyi, et al.
Published: (2025)
by: Cai, Zhongyi, et al.
Published: (2025)
Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence
by: Zheng, Yufei, et al.
Published: (2026)
by: Zheng, Yufei, et al.
Published: (2026)
Learning Adaptive Reasoning Paths for Efficient Visual Reasoning
by: Huang, Yixu, et al.
Published: (2026)
by: Huang, Yixu, et al.
Published: (2026)
Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
by: Jiang, Haichao, et al.
Published: (2026)
by: Jiang, Haichao, et al.
Published: (2026)
XRDSLAM: A Flexible and Modular Framework for Deep Learning based SLAM
by: Wang, Xiaomeng, et al.
Published: (2024)
by: Wang, Xiaomeng, et al.
Published: (2024)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
by: Yang, Qian, et al.
Published: (2026)
by: Yang, Qian, et al.
Published: (2026)
Rethinking Multi-domain Generalization with A General Learning Objective
by: Tan, Zhaorui, et al.
Published: (2024)
by: Tan, Zhaorui, et al.
Published: (2024)
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
by: Deng, Nianchen, et al.
Published: (2025)
by: Deng, Nianchen, et al.
Published: (2025)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
by: Cheng, An-Chieh, et al.
Published: (2024)
by: Cheng, An-Chieh, et al.
Published: (2024)
Revisiting Mutual Information Maximization for Generalized Category Discovery
by: Tan, Zhaorui, et al.
Published: (2024)
by: Tan, Zhaorui, et al.
Published: (2024)
Smooth Operator: Smooth Verifiable Reward Activates Spatial Reasoning Ability of Vision-Language Model
by: Jiao, Siwen, et al.
Published: (2026)
by: Jiao, Siwen, et al.
Published: (2026)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
by: Ma, Wenxin, et al.
Published: (2026)
by: Ma, Wenxin, et al.
Published: (2026)
Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
by: Zhu, Lanyun, et al.
Published: (2025)
by: Zhu, Lanyun, et al.
Published: (2025)
OSR-ViT: A Simple and Modular Framework for Open-Set Object Detection and Discovery
by: Inkawhich, Matthew, et al.
Published: (2024)
by: Inkawhich, Matthew, et al.
Published: (2024)
Real-Time Glottis Detection Framework via Spatial-decoupled Feature Learning for Nasal Transnasal Intubation
by: Liu, Jinyu, et al.
Published: (2026)
by: Liu, Jinyu, et al.
Published: (2026)
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
by: Zhao, Yu, et al.
Published: (2025)
by: Zhao, Yu, et al.
Published: (2025)
TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning
by: Wu, Nemin, et al.
Published: (2024)
by: Wu, Nemin, et al.
Published: (2024)
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
by: Qiu, Congpei, et al.
Published: (2026)
by: Qiu, Congpei, et al.
Published: (2026)
ElasticLaneNet: An Efficient Geometry-Flexible Approach for Lane Detection
by: Feng, Yaxin, et al.
Published: (2023)
by: Feng, Yaxin, et al.
Published: (2023)
Multimodal Graph Network Modeling for Human-Object Interaction Detection with PDE Graph Diffusion
by: Ji, Wenxuan, et al.
Published: (2025)
by: Ji, Wenxuan, et al.
Published: (2025)
KonfAI: A Modular and Fully Configurable Framework for Deep Learning in Medical Imaging
by: Boussot, Valentin, et al.
Published: (2025)
by: Boussot, Valentin, et al.
Published: (2025)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search
by: Liang, Tianming, et al.
Published: (2026)
by: Liang, Tianming, et al.
Published: (2026)
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
GaitKD: A Universal Decoupled Distillation Framework for Efficient Gait Recognition
by: Li, Yuqi, et al.
Published: (2026)
by: Li, Yuqi, et al.
Published: (2026)
Similar Items
-
GravCal: Single-Image Calibration of IMU Gravity Priors with Per-Sample Confidence
by: Zhu, Haichao, et al.
Published: (2026) -
Ultron: Enabling Temporal Geometry Compression of 3D Mesh Sequences using Temporal Correspondence and Mesh Deformation
by: Zhu, Haichao
Published: (2024) -
Sharpening Your Density Fields: Spiking Neuron Aided Fast Geometry Learning
by: Gu, Yi, et al.
Published: (2024) -
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
by: Li, Haoyuan, et al.
Published: (2026) -
SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
by: Guo, Jiajie, et al.
Published: (2025)