GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Liangyu, Zhao, Yingxiu, Wang, Jingyun, Wang, Yingyao, Pi, Bu, Wang, Chen, Zhang, Mingliang, Gu, Jihao, Li, Xiang, Zhu, Xiaoyong, Song, Jun, Zheng, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
von: Liu, Ruiheng, et al.
Veröffentlicht: (2026)
von: Liu, Ruiheng, et al.
Veröffentlicht: (2026)
Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
von: Gu, Jihao, et al.
Veröffentlicht: (2024)
von: Gu, Jihao, et al.
Veröffentlicht: (2024)
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
von: Chen, Peng, et al.
Veröffentlicht: (2025)
von: Chen, Peng, et al.
Veröffentlicht: (2025)
GeoSense-AI: Fast Location Inference from Crisis Microblogs
von: Sapru, Deepit
Veröffentlicht: (2025)
von: Sapru, Deepit
Veröffentlicht: (2025)
InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
von: Ai, Qihang, et al.
Veröffentlicht: (2025)
von: Ai, Qihang, et al.
Veröffentlicht: (2025)
Identification of Multivariate Measurement Error Models
von: Hu, Yingyao
Veröffentlicht: (2025)
von: Hu, Yingyao
Veröffentlicht: (2025)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
von: Cao, Meng, et al.
Veröffentlicht: (2025)
von: Cao, Meng, et al.
Veröffentlicht: (2025)
AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
von: Cao, Yue, et al.
Veröffentlicht: (2025)
von: Cao, Yue, et al.
Veröffentlicht: (2025)
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
von: Li, Zejun, et al.
Veröffentlicht: (2025)
von: Li, Zejun, et al.
Veröffentlicht: (2025)
LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning
von: Xu, Haiying, et al.
Veröffentlicht: (2026)
von: Xu, Haiying, et al.
Veröffentlicht: (2026)
DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
von: Xu, Xinrun, et al.
Veröffentlicht: (2025)
von: Xu, Xinrun, et al.
Veröffentlicht: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
Geoparsing: Diagram Parsing for Plane and Solid Geometry with a Unified Formal Language
von: Wang, Peijie, et al.
Veröffentlicht: (2026)
von: Wang, Peijie, et al.
Veröffentlicht: (2026)
EagleVision: Object-level Attribute Multimodal LLM for Remote Sensing
von: Jiang, Hongxiang, et al.
Veröffentlicht: (2025)
von: Jiang, Hongxiang, et al.
Veröffentlicht: (2025)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
von: Deng, Chao, et al.
Veröffentlicht: (2024)
von: Deng, Chao, et al.
Veröffentlicht: (2024)
Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case
von: Chen, Peng, et al.
Veröffentlicht: (2024)
von: Chen, Peng, et al.
Veröffentlicht: (2024)
GeoPro-VO: Dynamic Obstacle Avoidance with Geometric Projector Based on Velocity Obstacle
von: Huang, Jihao, et al.
Veröffentlicht: (2024)
von: Huang, Jihao, et al.
Veröffentlicht: (2024)
Performance Analysis of Traditional VQA Models Under Limited Computational Resources
von: Gu, Jihao
Veröffentlicht: (2025)
von: Gu, Jihao
Veröffentlicht: (2025)
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning
von: Gao, Xianqiang, et al.
Veröffentlicht: (2026)
von: Gao, Xianqiang, et al.
Veröffentlicht: (2026)
RSHallu: Dual-Mode Hallucination Evaluation for Remote-Sensing Multimodal Large Language Models with Domain-Tailored Mitigation
von: Zhou, Zihui, et al.
Veröffentlicht: (2026)
von: Zhou, Zihui, et al.
Veröffentlicht: (2026)
NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation
von: Wu, Weiming, et al.
Veröffentlicht: (2025)
von: Wu, Weiming, et al.
Veröffentlicht: (2025)
EFLNet: Enhancing Feature Learning for Infrared Small Target Detection
von: Yang, Bo, et al.
Veröffentlicht: (2023)
von: Yang, Bo, et al.
Veröffentlicht: (2023)
GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning
von: Jing, Jinhao, et al.
Veröffentlicht: (2026)
von: Jing, Jinhao, et al.
Veröffentlicht: (2026)
MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
MR. Judge: Multimodal Reasoner as a Judge
von: Pi, Renjie, et al.
Veröffentlicht: (2025)
von: Pi, Renjie, et al.
Veröffentlicht: (2025)
A Review of Detection, Evolution, and Data Reconstruction Strategies for False Data Injection Attacks in Power Cyber-Physical Systems
von: Bo, Xiaoyong
Veröffentlicht: (2025)
von: Bo, Xiaoyong
Veröffentlicht: (2025)
Artificial Orca Optimiser: Theory and Applications for Global Optimisation Problems
von: Lin Wang, et al.
Veröffentlicht: (2025)
von: Lin Wang, et al.
Veröffentlicht: (2025)
Effects of Sanqi Shengyu External Application Cream and Pulsed Electromagnetic Field on Knee Osteoarthritis in Older Adults: A Randomized Controlled Trial
von: Ruirui Wang, et al.
Veröffentlicht: (2025)
von: Ruirui Wang, et al.
Veröffentlicht: (2025)
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
GeoSketch: A Neural-Symbolic Approach to Geometric Multimodal Reasoning with Auxiliary Line Construction and Affine Transformation
von: Weng, Shichao, et al.
Veröffentlicht: (2025)
von: Weng, Shichao, et al.
Veröffentlicht: (2025)
An Artificial Intelligence Approach for Test‐Free Identification of Sarcopenia
von: Liangyu Yin, et al.
Veröffentlicht: (2024)
von: Liangyu Yin, et al.
Veröffentlicht: (2024)
GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
von: Wang, Fengxiang, et al.
Veröffentlicht: (2025)
von: Wang, Fengxiang, et al.
Veröffentlicht: (2025)
Concise Geometric Description as a Bridge: Unleashing the Potential of LLM for Plane Geometry Problem Solving
von: Wang, Jingyun, et al.
Veröffentlicht: (2026)
von: Wang, Jingyun, et al.
Veröffentlicht: (2026)
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
von: Cai, Shihao, et al.
Veröffentlicht: (2024)
von: Cai, Shihao, et al.
Veröffentlicht: (2024)
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
von: Xu, Ruoran, et al.
Veröffentlicht: (2026)
von: Xu, Ruoran, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
von: Liu, Ruiheng, et al.
Veröffentlicht: (2026) -
Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
von: Gu, Jihao, et al.
Veröffentlicht: (2024) -
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
von: Gu, Jihao, et al.
Veröffentlicht: (2025) -
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
von: Gu, Jihao, et al.
Veröffentlicht: (2025) -
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
von: Chen, Peng, et al.
Veröffentlicht: (2025)