SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jianshu, Li, Yijiang, Chen, Huifeixin, Lu, Haoran, Xue, Letian, Wang, Bingyang, Liu, Han |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
FairReason: Balancing Reasoning and Social Bias in MLLMs
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
von: Yu, Songsong, et al.
Veröffentlicht: (2025)
von: Yu, Songsong, et al.
Veröffentlicht: (2025)
Vision Language Models Know Law of Conservation without Understanding More-or-Less
von: Luo, Dezhi, et al.
Veröffentlicht: (2024)
von: Luo, Dezhi, et al.
Veröffentlicht: (2024)
How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
von: Yang, Zhen, et al.
Veröffentlicht: (2026)
von: Yang, Zhen, et al.
Veröffentlicht: (2026)
3D Primitives are a Spatial Language for VLMs
von: Liu, Junze, et al.
Veröffentlicht: (2026)
von: Liu, Junze, et al.
Veröffentlicht: (2026)
Vision Language Models Cannot Reason About Physical Transformation
von: Luo, Dezhi, et al.
Veröffentlicht: (2026)
von: Luo, Dezhi, et al.
Veröffentlicht: (2026)
Revisiting Data Augmentation in Deep Reinforcement Learning
von: Hu, Jianshu, et al.
Veröffentlicht: (2024)
von: Hu, Jianshu, et al.
Veröffentlicht: (2024)
Egocentric Bias in Vision-Language Models
von: Wang, Maijunxian, et al.
Veröffentlicht: (2026)
von: Wang, Maijunxian, et al.
Veröffentlicht: (2026)
Vision Language Models See What You Want but not What You See
von: Gao, Qingying, et al.
Veröffentlicht: (2024)
von: Gao, Qingying, et al.
Veröffentlicht: (2024)
See, Symbolize, Act: Grounding VLMs with Spatial Representations for Better Gameplay
von: Baghel, Ashish, et al.
Veröffentlicht: (2026)
von: Baghel, Ashish, et al.
Veröffentlicht: (2026)
How Do VLAs Effectively Inherit from VLMs?
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2025)
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2025)
Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks
von: Li, Chenjun
Veröffentlicht: (2026)
von: Li, Chenjun
Veröffentlicht: (2026)
Deconstructing Spatial Complexity: Hierarchical Decomposition for LLM Spatial Reasoning
von: Wang, Yi, et al.
Veröffentlicht: (2026)
von: Wang, Yi, et al.
Veröffentlicht: (2026)
SplitAgent: A Privacy-Preserving Distributed Architecture for Enterprise-Cloud Agent Collaboration
von: She, Jianshu
Veröffentlicht: (2026)
von: She, Jianshu
Veröffentlicht: (2026)
GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis
von: Liu, Haoyang, et al.
Veröffentlicht: (2025)
von: Liu, Haoyang, et al.
Veröffentlicht: (2025)
The Philosophical Foundations of Growing AI Like A Child
von: Luo, Dezhi, et al.
Veröffentlicht: (2025)
von: Luo, Dezhi, et al.
Veröffentlicht: (2025)
MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form Understanding
von: Ma, Jiefeng, et al.
Veröffentlicht: (2024)
von: Ma, Jiefeng, et al.
Veröffentlicht: (2024)
Probing Mechanical Reasoning in Large Vision Language Models
von: Sun, Haoran, et al.
Veröffentlicht: (2024)
von: Sun, Haoran, et al.
Veröffentlicht: (2024)
AdvEvo-MARL: Shaping Internalized Safety through Adversarial Co-Evolution in Multi-Agent Reinforcement Learning
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
von: Chen, Tianyu, et al.
Veröffentlicht: (2025)
von: Chen, Tianyu, et al.
Veröffentlicht: (2025)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving
von: Lian, Weitong, et al.
Veröffentlicht: (2026)
von: Lian, Weitong, et al.
Veröffentlicht: (2026)
Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
Evaluating VLMs' Spatial Reasoning Over Robot Motion: A Step Towards Robot Planning with Motion Preferences
von: Wu, Wenxi, et al.
Veröffentlicht: (2026)
von: Wu, Wenxi, et al.
Veröffentlicht: (2026)
Core Knowledge Deficits in Multi-Modal Language Models
von: Li, Yijiang, et al.
Veröffentlicht: (2024)
von: Li, Yijiang, et al.
Veröffentlicht: (2024)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
von: Liu, Yue, et al.
Veröffentlicht: (2025)
von: Liu, Yue, et al.
Veröffentlicht: (2025)
COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
von: Xia, Canming, et al.
Veröffentlicht: (2026)
von: Xia, Canming, et al.
Veröffentlicht: (2026)
Revisiting, Benchmarking and Understanding Unsupervised Graph Domain Adaptation
von: Liu, Meihan, et al.
Veröffentlicht: (2024)
von: Liu, Meihan, et al.
Veröffentlicht: (2024)
ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization
von: Yang, Letian, et al.
Veröffentlicht: (2026)
von: Yang, Letian, et al.
Veröffentlicht: (2026)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
Authorize-on-Demand: Dynamic Authorization with Legality-Aware Intellectual Property Protection for VLMs
von: Wang, Lianyu, et al.
Veröffentlicht: (2026)
von: Wang, Lianyu, et al.
Veröffentlicht: (2026)
Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
von: Lu, Shuai, et al.
Veröffentlicht: (2026)
von: Lu, Shuai, et al.
Veröffentlicht: (2026)
PyFi: Toward Pyramid-like Financial Image Understanding for VLMs via Adversarial Agents
von: Zhang, Yuqun, et al.
Veröffentlicht: (2025)
von: Zhang, Yuqun, et al.
Veröffentlicht: (2025)
Towards Improving Interpretability of Language Model Generation through a Structured Knowledge Discovery Approach
von: Liu, Shuqi, et al.
Veröffentlicht: (2025)
von: Liu, Shuqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025) -
FairReason: Balancing Reasoning and Social Bias in MLLMs
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025) -
How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
von: Yu, Songsong, et al.
Veröffentlicht: (2025) -
Vision Language Models Know Law of Conservation without Understanding More-or-Less
von: Luo, Dezhi, et al.
Veröffentlicht: (2024) -
How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
von: Yang, Zhen, et al.
Veröffentlicht: (2026)