GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Liangyu, Zhao, Yingxiu, Wang, Jingyun, Wang, Yingyao, Pi, Bu, Wang, Chen, Zhang, Mingliang, Gu, Jihao, Li, Xiang, Zhu, Xiaoyong, Song, Jun, Zheng, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
by: Liu, Ruiheng, et al.
Published: (2026)
by: Liu, Ruiheng, et al.
Published: (2026)
Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
by: Gu, Jihao, et al.
Published: (2024)
by: Gu, Jihao, et al.
Published: (2024)
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
by: Chen, Peng, et al.
Published: (2025)
by: Chen, Peng, et al.
Published: (2025)
GeoSense-AI: Fast Location Inference from Crisis Microblogs
by: Sapru, Deepit
Published: (2025)
by: Sapru, Deepit
Published: (2025)
InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
by: Ai, Qihang, et al.
Published: (2025)
by: Ai, Qihang, et al.
Published: (2025)
Identification of Multivariate Measurement Error Models
by: Hu, Yingyao
Published: (2025)
by: Hu, Yingyao
Published: (2025)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
by: Xie, Yiping, et al.
Published: (2026)
by: Xie, Yiping, et al.
Published: (2026)
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
by: Cao, Yue, et al.
Published: (2025)
by: Cao, Yue, et al.
Published: (2025)
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
by: Li, Zejun, et al.
Published: (2025)
by: Li, Zejun, et al.
Published: (2025)
LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning
by: Xu, Haiying, et al.
Published: (2026)
by: Xu, Haiying, et al.
Published: (2026)
DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
by: Xu, Xinrun, et al.
Published: (2025)
by: Xu, Xinrun, et al.
Published: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
by: Hu, Pengfei, et al.
Published: (2025)
by: Hu, Pengfei, et al.
Published: (2025)
Geoparsing: Diagram Parsing for Plane and Solid Geometry with a Unified Formal Language
by: Wang, Peijie, et al.
Published: (2026)
by: Wang, Peijie, et al.
Published: (2026)
EagleVision: Object-level Attribute Multimodal LLM for Remote Sensing
by: Jiang, Hongxiang, et al.
Published: (2025)
by: Jiang, Hongxiang, et al.
Published: (2025)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
by: Deng, Chao, et al.
Published: (2024)
by: Deng, Chao, et al.
Published: (2024)
Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case
by: Chen, Peng, et al.
Published: (2024)
by: Chen, Peng, et al.
Published: (2024)
GeoPro-VO: Dynamic Obstacle Avoidance with Geometric Projector Based on Velocity Obstacle
by: Huang, Jihao, et al.
Published: (2024)
by: Huang, Jihao, et al.
Published: (2024)
Performance Analysis of Traditional VQA Models Under Limited Computational Resources
by: Gu, Jihao
Published: (2025)
by: Gu, Jihao
Published: (2025)
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
by: Li, Jingyao, et al.
Published: (2025)
by: Li, Jingyao, et al.
Published: (2025)
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning
by: Gao, Xianqiang, et al.
Published: (2026)
by: Gao, Xianqiang, et al.
Published: (2026)
RSHallu: Dual-Mode Hallucination Evaluation for Remote-Sensing Multimodal Large Language Models with Domain-Tailored Mitigation
by: Zhou, Zihui, et al.
Published: (2026)
by: Zhou, Zihui, et al.
Published: (2026)
NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation
by: Wu, Weiming, et al.
Published: (2025)
by: Wu, Weiming, et al.
Published: (2025)
EFLNet: Enhancing Feature Learning for Infrared Small Target Detection
by: Yang, Bo, et al.
Published: (2023)
by: Yang, Bo, et al.
Published: (2023)
GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning
by: Jing, Jinhao, et al.
Published: (2026)
by: Jing, Jinhao, et al.
Published: (2026)
MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
by: Feng, Yuan, et al.
Published: (2025)
by: Feng, Yuan, et al.
Published: (2025)
MR. Judge: Multimodal Reasoner as a Judge
by: Pi, Renjie, et al.
Published: (2025)
by: Pi, Renjie, et al.
Published: (2025)
A Review of Detection, Evolution, and Data Reconstruction Strategies for False Data Injection Attacks in Power Cyber-Physical Systems
by: Bo, Xiaoyong
Published: (2025)
by: Bo, Xiaoyong
Published: (2025)
Artificial Orca Optimiser: Theory and Applications for Global Optimisation Problems
by: Lin Wang, et al.
Published: (2025)
by: Lin Wang, et al.
Published: (2025)
Effects of Sanqi Shengyu External Application Cream and Pulsed Electromagnetic Field on Knee Osteoarthritis in Older Adults: A Randomized Controlled Trial
by: Ruirui Wang, et al.
Published: (2025)
by: Ruirui Wang, et al.
Published: (2025)
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
by: Xia, Yinan, et al.
Published: (2025)
by: Xia, Yinan, et al.
Published: (2025)
GeoSketch: A Neural-Symbolic Approach to Geometric Multimodal Reasoning with Auxiliary Line Construction and Affine Transformation
by: Weng, Shichao, et al.
Published: (2025)
by: Weng, Shichao, et al.
Published: (2025)
An Artificial Intelligence Approach for Test‐Free Identification of Sarcopenia
by: Liangyu Yin, et al.
Published: (2024)
by: Liangyu Yin, et al.
Published: (2024)
GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
Concise Geometric Description as a Bridge: Unleashing the Potential of LLM for Plane Geometry Problem Solving
by: Wang, Jingyun, et al.
Published: (2026)
by: Wang, Jingyun, et al.
Published: (2026)
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
by: Cai, Shihao, et al.
Published: (2024)
by: Cai, Shihao, et al.
Published: (2024)
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
by: Xu, Ruoran, et al.
Published: (2026)
by: Xu, Ruoran, et al.
Published: (2026)
Similar Items
-
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
by: Liu, Ruiheng, et al.
Published: (2026) -
Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
by: Gu, Jihao, et al.
Published: (2024) -
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
by: Gu, Jihao, et al.
Published: (2025) -
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
by: Gu, Jihao, et al.
Published: (2025) -
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
by: Chen, Peng, et al.
Published: (2025)