Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Shuo, Cheng, Jianjie, Xu, Yinuo, Yu, Yongcan, Sheng, Lijun, Wang, Peijie, Jiang, Siru, Hu, Yongguan, Ling, Run, Shao, Yihua, Ma, Ao, Feng, Wei, He, Lingxiao, Wang, Meng, Xie, Qianlong, Wang, Xingxing, Sebe, Nicu, He, Ran, Liang, Jian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
di: Yu, Yongcan, et al.
Pubblicazione: (2025)
di: Yu, Yongcan, et al.
Pubblicazione: (2025)
How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1
di: Xu, Yinuo, et al.
Pubblicazione: (2026)
di: Xu, Yinuo, et al.
Pubblicazione: (2026)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
di: Yu, Yongcan, et al.
Pubblicazione: (2026)
di: Yu, Yongcan, et al.
Pubblicazione: (2026)
FITRep: Attention-Guided Item Representation via MLLMs
di: Zhang, Guoxiao, et al.
Pubblicazione: (2025)
di: Zhang, Guoxiao, et al.
Pubblicazione: (2025)
DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing
di: Lu, Shuo, et al.
Pubblicazione: (2025)
di: Lu, Shuo, et al.
Pubblicazione: (2025)
Do MLLMs Really Understand the Charts?
di: Zhang, Xiao, et al.
Pubblicazione: (2025)
di: Zhang, Xiao, et al.
Pubblicazione: (2025)
WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis
di: Lu, Shuo, et al.
Pubblicazione: (2026)
di: Lu, Shuo, et al.
Pubblicazione: (2026)
STAMP: Outlier-Aware Test-Time Adaptation with Stable Memory Replay
di: Yu, Yongcan, et al.
Pubblicazione: (2024)
di: Yu, Yongcan, et al.
Pubblicazione: (2024)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
di: Tang, Hao, et al.
Pubblicazione: (2025)
di: Tang, Hao, et al.
Pubblicazione: (2025)
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
di: Tang, Hao, et al.
Pubblicazione: (2024)
di: Tang, Hao, et al.
Pubblicazione: (2024)
Cooperative Pseudo Labeling for Unsupervised Federated Classification
di: Guo, Kuangpu, et al.
Pubblicazione: (2025)
di: Guo, Kuangpu, et al.
Pubblicazione: (2025)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
di: Song, Yue, et al.
Pubblicazione: (2023)
di: Song, Yue, et al.
Pubblicazione: (2023)
High-Fidelity 3D Facial Avatar Synthesis with Controllable Fine-Grained Expressions
di: He, Yikang, et al.
Pubblicazione: (2026)
di: He, Yikang, et al.
Pubblicazione: (2026)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
di: Wang, Yanbo, et al.
Pubblicazione: (2025)
di: Wang, Yanbo, et al.
Pubblicazione: (2025)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
di: Yu, Yongcan, et al.
Pubblicazione: (2025)
di: Yu, Yongcan, et al.
Pubblicazione: (2025)
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
di: Wang, Yanbo, et al.
Pubblicazione: (2025)
di: Wang, Yanbo, et al.
Pubblicazione: (2025)
Asymmetric GANs for Image-to-Image Translation
di: Tang, Hao, et al.
Pubblicazione: (2019)
di: Tang, Hao, et al.
Pubblicazione: (2019)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
di: Tang, Hao, et al.
Pubblicazione: (2025)
di: Tang, Hao, et al.
Pubblicazione: (2025)
Out-of-Distribution Detection: A Task-Oriented Survey of Recent Advances
di: Lu, Shuo, et al.
Pubblicazione: (2024)
di: Lu, Shuo, et al.
Pubblicazione: (2024)
Do Efficient Transformers Really Save Computation?
di: Yang, Kai, et al.
Pubblicazione: (2024)
di: Yang, Kai, et al.
Pubblicazione: (2024)
Large Language Models for Multimodal Deformable Image Registration
di: Ma, Mingrui, et al.
Pubblicazione: (2024)
di: Ma, Mingrui, et al.
Pubblicazione: (2024)
Rethinking the Learning Paradigm for Facial Expression Recognition
di: Wang, Weijie, et al.
Pubblicazione: (2022)
di: Wang, Weijie, et al.
Pubblicazione: (2022)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
di: Ou, Siqu, et al.
Pubblicazione: (2026)
di: Ou, Siqu, et al.
Pubblicazione: (2026)
RIA: A Ranking-Infused Approach for Optimized listwise CTR Prediction
di: Zhang, Guoxiao, et al.
Pubblicazione: (2025)
di: Zhang, Guoxiao, et al.
Pubblicazione: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
di: Gao, Shida, et al.
Pubblicazione: (2025)
di: Gao, Shida, et al.
Pubblicazione: (2025)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
di: Xing, Songlong, et al.
Pubblicazione: (2025)
di: Xing, Songlong, et al.
Pubblicazione: (2025)
Hyperbolic Busemann Neural Networks
di: Chen, Ziheng, et al.
Pubblicazione: (2026)
di: Chen, Ziheng, et al.
Pubblicazione: (2026)
Comment on ‘Aquablation for benign prostatic hyperplasia: real‐world prostate size relevance and bleeding events across 6 years’
di: Shuo Lin, et al.
Pubblicazione: (2026)
di: Shuo Lin, et al.
Pubblicazione: (2026)
PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems
di: Wang, Weijie, et al.
Pubblicazione: (2026)
di: Wang, Weijie, et al.
Pubblicazione: (2026)
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
di: Wang, Yidan, et al.
Pubblicazione: (2025)
di: Wang, Yidan, et al.
Pubblicazione: (2025)
RMLR: Extending Multinomial Logistic Regression into General Geometries
di: Chen, Ziheng, et al.
Pubblicazione: (2024)
di: Chen, Ziheng, et al.
Pubblicazione: (2024)
Reverse Personalization
di: Kung, Han-Wei, et al.
Pubblicazione: (2025)
di: Kung, Han-Wei, et al.
Pubblicazione: (2025)
One-dimensional $\mathbb{Z}$-classified topological crystalline insulator under space-time inversion symmetry
di: Lin, Ling, et al.
Pubblicazione: (2024)
di: Lin, Ling, et al.
Pubblicazione: (2024)
FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy
di: He, Qian, et al.
Pubblicazione: (2026)
di: He, Qian, et al.
Pubblicazione: (2026)
StyMam: A Mamba-Based Generator for Artistic Style Transfer
di: Hong, Zhou, et al.
Pubblicazione: (2026)
di: Hong, Zhou, et al.
Pubblicazione: (2026)
R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
di: Sheng, Lijun, et al.
Pubblicazione: (2025)
di: Sheng, Lijun, et al.
Pubblicazione: (2025)
On the Mathematics of RNA Velocity II: Algorithmic Aspects
di: Li, Tiejun, et al.
Pubblicazione: (2023)
di: Li, Tiejun, et al.
Pubblicazione: (2023)
RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers
di: Cao, Ke, et al.
Pubblicazione: (2025)
di: Cao, Ke, et al.
Pubblicazione: (2025)
Anti-Forgetting Adaptation for Unsupervised Person Re-identification
di: Chen, Hao, et al.
Pubblicazione: (2024)
di: Chen, Hao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
di: Yu, Yongcan, et al.
Pubblicazione: (2025) -
How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1
di: Xu, Yinuo, et al.
Pubblicazione: (2026) -
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
di: Yu, Yongcan, et al.
Pubblicazione: (2026) -
FITRep: Attention-Guided Item Representation via MLLMs
di: Zhang, Guoxiao, et al.
Pubblicazione: (2025) -
DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing
di: Lu, Shuo, et al.
Pubblicazione: (2025)