GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Ruiheng, Hao, Haihong, Han, Mingfei, Gu, Xin, Zhang, Kecheng, Li, Changlin, Chang, Xiaojun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
di: Hao, Haihong, et al.
Pubblicazione: (2025)
di: Hao, Haihong, et al.
Pubblicazione: (2025)
Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
di: Zhang, Kecheng, et al.
Pubblicazione: (2026)
di: Zhang, Kecheng, et al.
Pubblicazione: (2026)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
di: Hao, Haihong, et al.
Pubblicazione: (2026)
di: Hao, Haihong, et al.
Pubblicazione: (2026)
GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning
di: Sun, Jiayin, et al.
Pubblicazione: (2026)
di: Sun, Jiayin, et al.
Pubblicazione: (2026)
Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
di: Han, Mingfei, et al.
Pubblicazione: (2026)
di: Han, Mingfei, et al.
Pubblicazione: (2026)
Self-Consistency as a Free Lunch: Reducing Hallucinations in Vision-Language Models via Self-Reflection
di: Han, Mingfei, et al.
Pubblicazione: (2025)
di: Han, Mingfei, et al.
Pubblicazione: (2025)
Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
di: Han, Mingfei, et al.
Pubblicazione: (2023)
di: Han, Mingfei, et al.
Pubblicazione: (2023)
LongVLM: Efficient Long Video Understanding via Large Language Models
di: Weng, Yuetian, et al.
Pubblicazione: (2024)
di: Weng, Yuetian, et al.
Pubblicazione: (2024)
Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection
di: Zhang, Yuchen, et al.
Pubblicazione: (2026)
di: Zhang, Yuchen, et al.
Pubblicazione: (2026)
GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning
di: Miao, Deshui, et al.
Pubblicazione: (2026)
di: Miao, Deshui, et al.
Pubblicazione: (2026)
Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation
di: Jin, Minghao, et al.
Pubblicazione: (2026)
di: Jin, Minghao, et al.
Pubblicazione: (2026)
GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning
di: Xu, Liangyu, et al.
Pubblicazione: (2025)
di: Xu, Liangyu, et al.
Pubblicazione: (2025)
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning
di: Xu, Haiying, et al.
Pubblicazione: (2026)
di: Xu, Haiying, et al.
Pubblicazione: (2026)
Perception-Oriented Video Frame Interpolation via Asymmetric Blending
di: Wu, Guangyang, et al.
Pubblicazione: (2024)
di: Wu, Guangyang, et al.
Pubblicazione: (2024)
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
di: Qin, Zheng, et al.
Pubblicazione: (2025)
di: Qin, Zheng, et al.
Pubblicazione: (2025)
Which Layer Causes Distribution Deviation? Entropy-Guided Adaptive Pruning for Diffusion and Flow Models
di: Li, Changlin, et al.
Pubblicazione: (2025)
di: Li, Changlin, et al.
Pubblicazione: (2025)
Learning Consistent Taxonomic Classification through Hierarchical Reasoning
di: Li, Zhenghong, et al.
Pubblicazione: (2026)
di: Li, Zhenghong, et al.
Pubblicazione: (2026)
Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
di: Du, Yifan, et al.
Pubblicazione: (2025)
di: Du, Yifan, et al.
Pubblicazione: (2025)
IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale Benchmark
di: Cao, Zhe, et al.
Pubblicazione: (2025)
di: Cao, Zhe, et al.
Pubblicazione: (2025)
Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
di: Li, Changlin, et al.
Pubblicazione: (2025)
di: Li, Changlin, et al.
Pubblicazione: (2025)
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
di: Jiang, Longtao, et al.
Pubblicazione: (2025)
di: Jiang, Longtao, et al.
Pubblicazione: (2025)
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
di: Cao, Meng, et al.
Pubblicazione: (2026)
di: Cao, Meng, et al.
Pubblicazione: (2026)
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
di: Dai, Tingjun, et al.
Pubblicazione: (2026)
di: Dai, Tingjun, et al.
Pubblicazione: (2026)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
di: Gou, Yunhao, et al.
Pubblicazione: (2025)
di: Gou, Yunhao, et al.
Pubblicazione: (2025)
GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
di: Feng, Yuan, et al.
Pubblicazione: (2025)
di: Feng, Yuan, et al.
Pubblicazione: (2025)
Referring Remote Sensing Image Segmentation via Bidirectional Alignment Guided Joint Prediction
di: Zhang, Tianxiang, et al.
Pubblicazione: (2025)
di: Zhang, Tianxiang, et al.
Pubblicazione: (2025)
SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning
di: Yang, Xiao, et al.
Pubblicazione: (2026)
di: Yang, Xiao, et al.
Pubblicazione: (2026)
RC-GeoCP: Geometric Consensus for Radar-Camera Collaborative Perception
di: Bai, Xiaokai, et al.
Pubblicazione: (2026)
di: Bai, Xiaokai, et al.
Pubblicazione: (2026)
GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning
di: Jing, Jinhao, et al.
Pubblicazione: (2026)
di: Jing, Jinhao, et al.
Pubblicazione: (2026)
Unifying UAV Cross-View Geo-Localization via 3D Geometric Perception
di: Li, Haoyuan, et al.
Pubblicazione: (2026)
di: Li, Haoyuan, et al.
Pubblicazione: (2026)
GeoFocus: Blending Efficient Global-to-Local Perception for Multimodal Geometry Problem-Solving
di: Deng, Linger, et al.
Pubblicazione: (2026)
di: Deng, Linger, et al.
Pubblicazione: (2026)
Mitigating Data Redundancy to Revitalize Transformer-based Long-Term Time Series Forecasting System
di: Li, Mingjie, et al.
Pubblicazione: (2022)
di: Li, Mingjie, et al.
Pubblicazione: (2022)
Efficient Training of Large Vision Models via Advanced Automated Progressive Learning
di: Li, Changlin, et al.
Pubblicazione: (2024)
di: Li, Changlin, et al.
Pubblicazione: (2024)
Vision-Centric Activation and Coordination for Multimodal Large Language Models
di: Wang, Yunnan, et al.
Pubblicazione: (2025)
di: Wang, Yunnan, et al.
Pubblicazione: (2025)
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding
di: Zhu, Jiashun, et al.
Pubblicazione: (2026)
di: Zhu, Jiashun, et al.
Pubblicazione: (2026)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
di: Xiao, Tong, et al.
Pubblicazione: (2025)
di: Xiao, Tong, et al.
Pubblicazione: (2025)
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
di: Lu, Jinda, et al.
Pubblicazione: (2026)
di: Lu, Jinda, et al.
Pubblicazione: (2026)
Perception-Aware Multimodal Spatial Reasoning from Monocular Images
di: Cheng, Yanchun, et al.
Pubblicazione: (2026)
di: Cheng, Yanchun, et al.
Pubblicazione: (2026)
Documenti analoghi
-
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
di: Hao, Haihong, et al.
Pubblicazione: (2025) -
Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
di: Zhang, Kecheng, et al.
Pubblicazione: (2026) -
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
di: Hao, Haihong, et al.
Pubblicazione: (2026) -
GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning
di: Sun, Jiayin, et al.
Pubblicazione: (2026) -
Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
di: Han, Mingfei, et al.
Pubblicazione: (2026)