Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Shan, Chen, Aotian, Sun, Yanpeng, Gu, Jindong, Zheng, Yi-Yu, Koniusz, Piotr, Zou, Kai, Hengel, Anton van den, Xue, Yuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
by: Sun, Yanpeng, et al.
Published: (2025)
by: Sun, Yanpeng, et al.
Published: (2025)
Hierarchical Process Reward Models are Symbolic Vision Learners
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
by: Li, Xiaoyuan, et al.
Published: (2025)
by: Li, Xiaoyuan, et al.
Published: (2025)
Artemis: Structured Visual Reasoning for Perception Policy Learning
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
Learning Gaussian Representation for Eye Fixation Prediction
by: Song, Peipei, et al.
Published: (2024)
by: Song, Peipei, et al.
Published: (2024)
PACE: Marrying generalization in PArameter-efficient fine-tuning with Consistency rEgularization
by: Ni, Yao, et al.
Published: (2024)
by: Ni, Yao, et al.
Published: (2024)
Video Understanding by Design: How Datasets Shape Architectures and Insights
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
OpenKD: Opening Prompt Diversity for Zero- and Few-shot Keypoint Detection
by: Lu, Changsheng, et al.
Published: (2024)
by: Lu, Changsheng, et al.
Published: (2024)
Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
MuseBarControl: Enhancing Fine-Grained Control in Symbolic Music Generation through Pre-Training and Counterfactual Loss
by: Shu, Yangyang, et al.
Published: (2024)
by: Shu, Yangyang, et al.
Published: (2024)
Frame-wise Conditioning Adaptation for Fine-Tuning Diffusion Models in Text-to-Video Prediction
by: Liu, Zheyuan, et al.
Published: (2025)
by: Liu, Zheyuan, et al.
Published: (2025)
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
by: Albert, Paul, et al.
Published: (2025)
by: Albert, Paul, et al.
Published: (2025)
CHAIN: Enhancing Generalization in Data-Efficient GANs via lipsCHitz continuity constrAIned Normalization
by: Ni, Yao, et al.
Published: (2024)
by: Ni, Yao, et al.
Published: (2024)
Feature Hallucination for Self-supervised Action Recognition
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
by: Lu, Shuo, et al.
Published: (2026)
by: Lu, Shuo, et al.
Published: (2026)
Confident Sinkhorn Allocation for Pseudo-Labeling
by: Nguyen, Vu, et al.
Published: (2022)
by: Nguyen, Vu, et al.
Published: (2022)
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
by: Xia, Jiatong, et al.
Published: (2026)
by: Xia, Jiatong, et al.
Published: (2026)
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems
by: Chen, Shuhang, et al.
Published: (2025)
by: Chen, Shuhang, et al.
Published: (2025)
Chem4DLLM: 4D Multimodal LLMs for Chemical Dynamics Understanding
by: Li, Xinyu, et al.
Published: (2026)
by: Li, Xinyu, et al.
Published: (2026)
Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack
by: Gu, Chenhe, et al.
Published: (2025)
by: Gu, Chenhe, et al.
Published: (2025)
Uncertainty-DTW for Sequences and Visual Tokens
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
by: Qiang, Chenhui, et al.
Published: (2025)
by: Qiang, Chenhui, et al.
Published: (2025)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
by: Huang, Jincai, et al.
Published: (2026)
by: Huang, Jincai, et al.
Published: (2026)
DermoGPT: Open Weights and Open Data for Morphology-Grounded Dermatological Reasoning MLLMs
by: Ru, Jinghan, et al.
Published: (2026)
by: Ru, Jinghan, et al.
Published: (2026)
Graph Self-Supervised Learning with Learnable Structural and Positional Encodings
by: Wijesinghe, Asiri, et al.
Published: (2025)
by: Wijesinghe, Asiri, et al.
Published: (2025)
Let Your Video Listen to Your Music!
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Procedural Pretraining: Warming Up Language Models with Abstract Data
by: Jiang, Liangze, et al.
Published: (2026)
by: Jiang, Liangze, et al.
Published: (2026)
MARec: Metadata Alignment for cold-start Recommendation
by: Monteil, Julien, et al.
Published: (2024)
by: Monteil, Julien, et al.
Published: (2024)
Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings
by: Yin, Wei, et al.
Published: (2022)
by: Yin, Wei, et al.
Published: (2022)
Premonition: Using Generative Models to Preempt Future Data Changes in Continual Learning
by: McDonnell, Mark D., et al.
Published: (2024)
by: McDonnell, Mark D., et al.
Published: (2024)
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
by: Xiao, Rui, et al.
Published: (2026)
by: Xiao, Rui, et al.
Published: (2026)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
by: He, Wei, et al.
Published: (2024)
by: He, Wei, et al.
Published: (2024)
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
by: Wu, Yixuan, et al.
Published: (2025)
by: Wu, Yixuan, et al.
Published: (2025)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
by: Tong, Jintao, et al.
Published: (2025)
by: Tong, Jintao, et al.
Published: (2025)
Continual Learning on CLIP via Incremental Prompt Tuning with Intrinsic Textual Anchors
by: Lu, Haodong, et al.
Published: (2025)
by: Lu, Haodong, et al.
Published: (2025)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
by: Hwang, Hyeonbin, et al.
Published: (2024)
by: Hwang, Hyeonbin, et al.
Published: (2024)
Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs
by: Liu, Aoming, et al.
Published: (2025)
by: Liu, Aoming, et al.
Published: (2025)
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
by: Yin, Zhihan, et al.
Published: (2026)
by: Yin, Zhihan, et al.
Published: (2026)
Similar Items
-
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
by: Sun, Yanpeng, et al.
Published: (2025) -
Hierarchical Process Reward Models are Symbolic Vision Learners
by: Zhang, Shan, et al.
Published: (2025) -
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
by: Li, Xiaoyuan, et al.
Published: (2025) -
Artemis: Structured Visual Reasoning for Perception Policy Learning
by: Tang, Wei, et al.
Published: (2025) -
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
by: Zhu, Jie, et al.
Published: (2026)