MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
Fuente:
arXiv
Salvato in:
| Autori principali: | Liang, Hao, Sun, Linzhuang, Zhou, Minxuan, Chen, Zirong, Qiang, Meiyi, Lin, Mingan, Li, Tianpeng, Yang, Fan, Zhou, Zenan, Zhang, Wentao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
di: Sun, Linzhuang, et al.
Pubblicazione: (2025)
di: Sun, Linzhuang, et al.
Pubblicazione: (2025)
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
di: Liang, Hao, et al.
Pubblicazione: (2025)
di: Liang, Hao, et al.
Pubblicazione: (2025)
BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search
di: Sun, Linzhuang, et al.
Pubblicazione: (2024)
di: Sun, Linzhuang, et al.
Pubblicazione: (2024)
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
di: Liu, Wentao, et al.
Pubblicazione: (2024)
di: Liu, Wentao, et al.
Pubblicazione: (2024)
BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
di: Liang, Hao, et al.
Pubblicazione: (2026)
di: Liang, Hao, et al.
Pubblicazione: (2026)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
di: Cai, Qifeng, et al.
Pubblicazione: (2025)
di: Cai, Qifeng, et al.
Pubblicazione: (2025)
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
di: Feng, Jun, et al.
Pubblicazione: (2025)
di: Feng, Jun, et al.
Pubblicazione: (2025)
Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
di: Feng, Hengyi, et al.
Pubblicazione: (2025)
di: Feng, Hengyi, et al.
Pubblicazione: (2025)
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection
di: Yan, Yibo, et al.
Pubblicazione: (2025)
di: Yan, Yibo, et al.
Pubblicazione: (2025)
Data Proportion Detection for Optimized Data Management for Large Language Models
di: Liang, Hao, et al.
Pubblicazione: (2024)
di: Liang, Hao, et al.
Pubblicazione: (2024)
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
di: Zhang, Tao, et al.
Pubblicazione: (2024)
di: Zhang, Tao, et al.
Pubblicazione: (2024)
Let's Verify Math Questions Step by Step
di: Shen, Chengyu, et al.
Pubblicazione: (2025)
di: Shen, Chengyu, et al.
Pubblicazione: (2025)
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
di: Xue, Boyang, et al.
Pubblicazione: (2025)
di: Xue, Boyang, et al.
Pubblicazione: (2025)
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
di: Chen, Mingyang, et al.
Pubblicazione: (2024)
di: Chen, Mingyang, et al.
Pubblicazione: (2024)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
di: Shi, Wenhao, et al.
Pubblicazione: (2024)
di: Shi, Wenhao, et al.
Pubblicazione: (2024)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
di: Wang, Lei, et al.
Pubblicazione: (2024)
di: Wang, Lei, et al.
Pubblicazione: (2024)
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
di: Wang, Yiming, et al.
Pubblicazione: (2025)
di: Wang, Yiming, et al.
Pubblicazione: (2025)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
di: Fang, Meng, et al.
Pubblicazione: (2024)
di: Fang, Meng, et al.
Pubblicazione: (2024)
ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
di: Wu, Yanan, et al.
Pubblicazione: (2024)
di: Wu, Yanan, et al.
Pubblicazione: (2024)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
di: Liang, Hao, et al.
Pubblicazione: (2024)
di: Liang, Hao, et al.
Pubblicazione: (2024)
BaichuanSEED: Sharing the Potential of ExtensivE Data Collection and Deduplication by Introducing a Competitive Large Language Model Baseline
di: Dong, Guosheng, et al.
Pubblicazione: (2024)
di: Dong, Guosheng, et al.
Pubblicazione: (2024)
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
di: Guo, Zile, et al.
Pubblicazione: (2026)
di: Guo, Zile, et al.
Pubblicazione: (2026)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
di: Shi, Weikang, et al.
Pubblicazione: (2025)
di: Shi, Weikang, et al.
Pubblicazione: (2025)
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
di: Colle, Vincenzo, et al.
Pubblicazione: (2025)
di: Colle, Vincenzo, et al.
Pubblicazione: (2025)
Auto311: A Confidence-guided Automated System for Non-emergency Calls
di: Chen, Zirong, et al.
Pubblicazione: (2023)
di: Chen, Zirong, et al.
Pubblicazione: (2023)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
di: Chen, Mingyang, et al.
Pubblicazione: (2025)
di: Chen, Mingyang, et al.
Pubblicazione: (2025)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
di: Zeng, Liang, et al.
Pubblicazione: (2024)
di: Zeng, Liang, et al.
Pubblicazione: (2024)
FineMath: A Fine-Grained Mathematical Evaluation Benchmark for Chinese Large Language Models
di: Liu, Yan, et al.
Pubblicazione: (2024)
di: Liu, Yan, et al.
Pubblicazione: (2024)
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
di: Wang, Ke, et al.
Pubblicazione: (2025)
di: Wang, Ke, et al.
Pubblicazione: (2025)
JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models
di: Hao, Yifan, et al.
Pubblicazione: (2025)
di: Hao, Yifan, et al.
Pubblicazione: (2025)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
di: Lu, Pan, et al.
Pubblicazione: (2023)
di: Lu, Pan, et al.
Pubblicazione: (2023)
Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models
di: Zhang, YiFan, et al.
Pubblicazione: (2024)
di: Zhang, YiFan, et al.
Pubblicazione: (2024)
VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing
di: Deng, Andong, et al.
Pubblicazione: (2026)
di: Deng, Andong, et al.
Pubblicazione: (2026)
PAS: Data-Efficient Plug-and-Play Prompt Augmentation System
di: Zheng, Miao, et al.
Pubblicazione: (2024)
di: Zheng, Miao, et al.
Pubblicazione: (2024)
DOP: Diagnostic-Oriented Prompting for Large Language Models in Mathematical Correction
di: Chen, Hao, et al.
Pubblicazione: (2024)
di: Chen, Hao, et al.
Pubblicazione: (2024)
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
di: Qiao, Runqi, et al.
Pubblicazione: (2024)
di: Qiao, Runqi, et al.
Pubblicazione: (2024)
MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation
di: Xue, Haochen, et al.
Pubblicazione: (2025)
di: Xue, Haochen, et al.
Pubblicazione: (2025)
LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE
di: Shang, Yu, et al.
Pubblicazione: (2025)
di: Shang, Yu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
di: Sun, Linzhuang, et al.
Pubblicazione: (2025) -
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
di: Liang, Hao, et al.
Pubblicazione: (2025) -
BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search
di: Sun, Linzhuang, et al.
Pubblicazione: (2024) -
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
di: Liu, Wentao, et al.
Pubblicazione: (2024) -
BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
di: Guo, Tianyu, et al.
Pubblicazione: (2025)