Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Shuo, Cheng, Jianjie, Xu, Yinuo, Yu, Yongcan, Sheng, Lijun, Wang, Peijie, Jiang, Siru, Hu, Yongguan, Ling, Run, Shao, Yihua, Ma, Ao, Feng, Wei, He, Lingxiao, Wang, Meng, Xie, Qianlong, Wang, Xingxing, Sebe, Nicu, He, Ran, Liang, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
von: Yu, Yongcan, et al.
Veröffentlicht: (2026)
von: Yu, Yongcan, et al.
Veröffentlicht: (2026)
FITRep: Attention-Guided Item Representation via MLLMs
von: Zhang, Guoxiao, et al.
Veröffentlicht: (2025)
von: Zhang, Guoxiao, et al.
Veröffentlicht: (2025)
DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing
von: Lu, Shuo, et al.
Veröffentlicht: (2025)
von: Lu, Shuo, et al.
Veröffentlicht: (2025)
Do MLLMs Really Understand the Charts?
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis
von: Lu, Shuo, et al.
Veröffentlicht: (2026)
von: Lu, Shuo, et al.
Veröffentlicht: (2026)
STAMP: Outlier-Aware Test-Time Adaptation with Stable Memory Replay
von: Yu, Yongcan, et al.
Veröffentlicht: (2024)
von: Yu, Yongcan, et al.
Veröffentlicht: (2024)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
Cooperative Pseudo Labeling for Unsupervised Federated Classification
von: Guo, Kuangpu, et al.
Veröffentlicht: (2025)
von: Guo, Kuangpu, et al.
Veröffentlicht: (2025)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
von: Song, Yue, et al.
Veröffentlicht: (2023)
von: Song, Yue, et al.
Veröffentlicht: (2023)
High-Fidelity 3D Facial Avatar Synthesis with Controllable Fine-Grained Expressions
von: He, Yikang, et al.
Veröffentlicht: (2026)
von: He, Yikang, et al.
Veröffentlicht: (2026)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
Asymmetric GANs for Image-to-Image Translation
von: Tang, Hao, et al.
Veröffentlicht: (2019)
von: Tang, Hao, et al.
Veröffentlicht: (2019)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
Out-of-Distribution Detection: A Task-Oriented Survey of Recent Advances
von: Lu, Shuo, et al.
Veröffentlicht: (2024)
von: Lu, Shuo, et al.
Veröffentlicht: (2024)
Do Efficient Transformers Really Save Computation?
von: Yang, Kai, et al.
Veröffentlicht: (2024)
von: Yang, Kai, et al.
Veröffentlicht: (2024)
Large Language Models for Multimodal Deformable Image Registration
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
Rethinking the Learning Paradigm for Facial Expression Recognition
von: Wang, Weijie, et al.
Veröffentlicht: (2022)
von: Wang, Weijie, et al.
Veröffentlicht: (2022)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
RIA: A Ranking-Infused Approach for Optimized listwise CTR Prediction
von: Zhang, Guoxiao, et al.
Veröffentlicht: (2025)
von: Zhang, Guoxiao, et al.
Veröffentlicht: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
von: Gao, Shida, et al.
Veröffentlicht: (2025)
von: Gao, Shida, et al.
Veröffentlicht: (2025)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
von: Xing, Songlong, et al.
Veröffentlicht: (2025)
von: Xing, Songlong, et al.
Veröffentlicht: (2025)
Hyperbolic Busemann Neural Networks
von: Chen, Ziheng, et al.
Veröffentlicht: (2026)
von: Chen, Ziheng, et al.
Veröffentlicht: (2026)
Comment on ‘Aquablation for benign prostatic hyperplasia: real‐world prostate size relevance and bleeding events across 6 years’
von: Shuo Lin, et al.
Veröffentlicht: (2026)
von: Shuo Lin, et al.
Veröffentlicht: (2026)
PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems
von: Wang, Weijie, et al.
Veröffentlicht: (2026)
von: Wang, Weijie, et al.
Veröffentlicht: (2026)
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
RMLR: Extending Multinomial Logistic Regression into General Geometries
von: Chen, Ziheng, et al.
Veröffentlicht: (2024)
von: Chen, Ziheng, et al.
Veröffentlicht: (2024)
Reverse Personalization
von: Kung, Han-Wei, et al.
Veröffentlicht: (2025)
von: Kung, Han-Wei, et al.
Veröffentlicht: (2025)
One-dimensional $\mathbb{Z}$-classified topological crystalline insulator under space-time inversion symmetry
von: Lin, Ling, et al.
Veröffentlicht: (2024)
von: Lin, Ling, et al.
Veröffentlicht: (2024)
FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy
von: He, Qian, et al.
Veröffentlicht: (2026)
von: He, Qian, et al.
Veröffentlicht: (2026)
StyMam: A Mamba-Based Generator for Artistic Style Transfer
von: Hong, Zhou, et al.
Veröffentlicht: (2026)
von: Hong, Zhou, et al.
Veröffentlicht: (2026)
R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
von: Sheng, Lijun, et al.
Veröffentlicht: (2025)
von: Sheng, Lijun, et al.
Veröffentlicht: (2025)
On the Mathematics of RNA Velocity II: Algorithmic Aspects
von: Li, Tiejun, et al.
Veröffentlicht: (2023)
von: Li, Tiejun, et al.
Veröffentlicht: (2023)
RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers
von: Cao, Ke, et al.
Veröffentlicht: (2025)
von: Cao, Ke, et al.
Veröffentlicht: (2025)
Anti-Forgetting Adaptation for Unsupervised Person Re-identification
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
von: Yu, Yongcan, et al.
Veröffentlicht: (2025) -
How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1
von: Xu, Yinuo, et al.
Veröffentlicht: (2026) -
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
von: Yu, Yongcan, et al.
Veröffentlicht: (2026) -
FITRep: Attention-Guided Item Representation via MLLMs
von: Zhang, Guoxiao, et al.
Veröffentlicht: (2025) -
DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing
von: Lu, Shuo, et al.
Veröffentlicht: (2025)