MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Renrui, Jiang, Dongzhi, Zhang, Yichi, Lin, Haokun, Guo, Ziyu, Qiu, Pengshuo, Zhou, Aojun, Lu, Pan, Chang, Kai-Wei, Gao, Peng, Li, Hongsheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
di: Lu, Zimu, et al.
Pubblicazione: (2024)
di: Lu, Zimu, et al.
Pubblicazione: (2024)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
di: Chen, Xinyan, et al.
Pubblicazione: (2025)
di: Chen, Xinyan, et al.
Pubblicazione: (2025)
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
di: Zhang, Boning, et al.
Pubblicazione: (2024)
di: Zhang, Boning, et al.
Pubblicazione: (2024)
MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
di: Shi, Weikang, et al.
Pubblicazione: (2025)
di: Shi, Weikang, et al.
Pubblicazione: (2025)
MathChat: Converse to Tackle Challenging Math Problems with LLM Agents
di: Wu, Yiran, et al.
Pubblicazione: (2023)
di: Wu, Yiran, et al.
Pubblicazione: (2023)
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
di: Lu, Zimu, et al.
Pubblicazione: (2024)
di: Lu, Zimu, et al.
Pubblicazione: (2024)
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
di: Wang, Ke, et al.
Pubblicazione: (2025)
di: Wang, Ke, et al.
Pubblicazione: (2025)
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
di: Guo, Dadi, et al.
Pubblicazione: (2026)
di: Guo, Dadi, et al.
Pubblicazione: (2026)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
di: Sun, Yanpeng, et al.
Pubblicazione: (2025)
di: Sun, Yanpeng, et al.
Pubblicazione: (2025)
MathVC: An LLM-Simulated Multi-Character Virtual Classroom for Mathematics Education
di: Yue, Murong, et al.
Pubblicazione: (2024)
di: Yue, Murong, et al.
Pubblicazione: (2024)
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
di: Qiao, Runqi, et al.
Pubblicazione: (2024)
di: Qiao, Runqi, et al.
Pubblicazione: (2024)
GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving
di: Zhang, Yingji, et al.
Pubblicazione: (2026)
di: Zhang, Yingji, et al.
Pubblicazione: (2026)
Solving Formal Math Problems by Decomposition and Iterative Reflection
di: Zhou, Yichi, et al.
Pubblicazione: (2025)
di: Zhou, Yichi, et al.
Pubblicazione: (2025)
VISTA: Visual Integrated System for Tailored Automation in Math Problem Generation Using LLM
di: Lee, Jeongwoo, et al.
Pubblicazione: (2024)
di: Lee, Jeongwoo, et al.
Pubblicazione: (2024)
SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
di: Lu, Xudong, et al.
Pubblicazione: (2024)
di: Lu, Xudong, et al.
Pubblicazione: (2024)
PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models
di: Guo, Zilu, et al.
Pubblicazione: (2025)
di: Guo, Zilu, et al.
Pubblicazione: (2025)
Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions
di: Li, Hang, et al.
Pubblicazione: (2024)
di: Li, Hang, et al.
Pubblicazione: (2024)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
di: Ying, Huaiyuan, et al.
Pubblicazione: (2024)
di: Ying, Huaiyuan, et al.
Pubblicazione: (2024)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
di: Fang, Meng, et al.
Pubblicazione: (2024)
di: Fang, Meng, et al.
Pubblicazione: (2024)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
di: Zhao, Yilun, et al.
Pubblicazione: (2023)
di: Zhao, Yilun, et al.
Pubblicazione: (2023)
World Models for Math Story Problems
di: Opedal, Andreas, et al.
Pubblicazione: (2023)
di: Opedal, Andreas, et al.
Pubblicazione: (2023)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
di: Qin, Tian, et al.
Pubblicazione: (2025)
di: Qin, Tian, et al.
Pubblicazione: (2025)
MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion
di: Pei, Qizhi, et al.
Pubblicazione: (2025)
di: Pei, Qizhi, et al.
Pubblicazione: (2025)
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
di: Xu, Yifan, et al.
Pubblicazione: (2024)
di: Xu, Yifan, et al.
Pubblicazione: (2024)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
di: Anantheswaran, Ujjwala, et al.
Pubblicazione: (2024)
di: Anantheswaran, Ujjwala, et al.
Pubblicazione: (2024)
SafeMath: Inference-time Safety improves Math Accuracy
di: Basu, Sagnik, et al.
Pubblicazione: (2026)
di: Basu, Sagnik, et al.
Pubblicazione: (2026)
SmartBench: Is Your LLM Truly a Good Chinese Smartphone Assistant?
di: Lu, Xudong, et al.
Pubblicazione: (2025)
di: Lu, Xudong, et al.
Pubblicazione: (2025)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
di: Huan, Maggie, et al.
Pubblicazione: (2025)
di: Huan, Maggie, et al.
Pubblicazione: (2025)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
di: Lu, Pan, et al.
Pubblicazione: (2023)
di: Lu, Pan, et al.
Pubblicazione: (2023)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
Adversarial Math Word Problem Generation
di: Xie, Roy, et al.
Pubblicazione: (2024)
di: Xie, Roy, et al.
Pubblicazione: (2024)
Does Machine Unlearning Truly Remove Knowledge?
di: Chen, Haokun, et al.
Pubblicazione: (2025)
di: Chen, Haokun, et al.
Pubblicazione: (2025)
mAceReason-Math: A Dataset of High-Quality Multilingual Math Problems Ready For RLVR
di: Dobler, Konstantin, et al.
Pubblicazione: (2026)
di: Dobler, Konstantin, et al.
Pubblicazione: (2026)
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems
di: Chen, Shuhang, et al.
Pubblicazione: (2025)
di: Chen, Shuhang, et al.
Pubblicazione: (2025)
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
di: Liu, Xianyang, et al.
Pubblicazione: (2025)
di: Liu, Xianyang, et al.
Pubblicazione: (2025)
MathBuddy: A Multimodal System for Affective Math Tutoring
di: Kar, Debanjana, et al.
Pubblicazione: (2025)
di: Kar, Debanjana, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
di: Guo, Ziyu, et al.
Pubblicazione: (2025) -
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
di: Lu, Zimu, et al.
Pubblicazione: (2024) -
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
di: Chen, Xinyan, et al.
Pubblicazione: (2025) -
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024) -
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
di: Zhang, Boning, et al.
Pubblicazione: (2024)