ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Yibo, Wang, Shen, Huo, Jiahao, Li, Hang, Li, Boyan, Su, Jiamin, Gao, Xiong, Zhang, Yi-Fan, Xu, Tianlong, Chu, Zhendong, Zhong, Aoxiao, Wang, Kun, Xiong, Hui, Yu, Philip S., Hu, Xuming, Wen, Qingsong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection
by: Yan, Yibo, et al.
Published: (2025)
by: Yan, Yibo, et al.
Published: (2025)
Multimodal AI Teacher: Integrating Edge Computing and Reasoning Models for Enhanced Student Error Analysis
by: Tianlong Xu, et al.
Published: (2025)
by: Tianlong Xu, et al.
Published: (2025)
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
by: Yan, Yibo, et al.
Published: (2024)
by: Yan, Yibo, et al.
Published: (2024)
AI-Driven Virtual Teacher for Enhanced Educational Efficiency: Leveraging Large Pretrain Models for Autonomous Error Analysis and Correction
by: Xu, Tianlong, et al.
Published: (2024)
by: Xu, Tianlong, et al.
Published: (2024)
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
by: Yan, Yibo, et al.
Published: (2025)
by: Yan, Yibo, et al.
Published: (2025)
LLM Agents for Education: Advances and Applications
by: Chu, Zhendong, et al.
Published: (2025)
by: Chu, Zhendong, et al.
Published: (2025)
MINER: Mining the Underlying Pattern of Modality-Specific Neurons in Multimodal Large Language Models
by: Huang, Kaichen, et al.
Published: (2024)
by: Huang, Kaichen, et al.
Published: (2024)
Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
by: Song, Dingjie, et al.
Published: (2026)
by: Song, Dingjie, et al.
Published: (2026)
From Correctness to Comprehension: AI Agents for Personalized Error Diagnosis in Education
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
by: Su, Jiamin, et al.
Published: (2025)
by: Su, Jiamin, et al.
Published: (2025)
Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions
by: Li, Hang, et al.
Published: (2024)
by: Li, Hang, et al.
Published: (2024)
MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model
by: Huo, Jiahao, et al.
Published: (2024)
by: Huo, Jiahao, et al.
Published: (2024)
Mind Scramble: Unveiling Large Language Model Psychology Via Typoglycemia
by: Yu, Miao, et al.
Published: (2024)
by: Yu, Miao, et al.
Published: (2024)
ARM2: Adaptive Reasoning Model with Vision Understanding and Executable Code
by: Xie, Jian, et al.
Published: (2025)
by: Xie, Jian, et al.
Published: (2025)
Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis
by: Huang, Haoming, et al.
Published: (2025)
by: Huang, Haoming, et al.
Published: (2025)
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
by: Huo, Jiahao, et al.
Published: (2025)
by: Huo, Jiahao, et al.
Published: (2025)
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey
by: Dang, Yunkai, et al.
Published: (2024)
by: Dang, Yunkai, et al.
Published: (2024)
NL2SQL-BUGs: A Benchmark for Detecting Semantic Errors in NL2SQL Translation
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
by: Yu, Erxin, et al.
Published: (2025)
by: Yu, Erxin, et al.
Published: (2025)
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
by: Zhang, Jianghangfan, et al.
Published: (2025)
by: Zhang, Jianghangfan, et al.
Published: (2025)
Automate Knowledge Concept Tagging on Math Questions with LLMs
by: Li, Hang, et al.
Published: (2024)
by: Li, Hang, et al.
Published: (2024)
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever
by: Li, Hang, et al.
Published: (2024)
by: Li, Hang, et al.
Published: (2024)
Knowledge Tagging with Large Language Model based Multi-Agent System
by: Li, Hang, et al.
Published: (2024)
by: Li, Hang, et al.
Published: (2024)
On the Relation Between LP Sharpness and Limiting Error Ratio and Complexity Implications for Restarted PDHG
by: Xiong, Zikai, et al.
Published: (2023)
by: Xiong, Zikai, et al.
Published: (2023)
SLMFix: Leveraging Small Language Models for Error Fixing with Reinforcement Learning
by: Fu, David Jiahao, et al.
Published: (2025)
by: Fu, David Jiahao, et al.
Published: (2025)
CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring
by: Su, Jiamin, et al.
Published: (2025)
by: Su, Jiamin, et al.
Published: (2025)
An Efficient Gradient-Aware Error-Bounded Lossy Compressor for Federated Learning
by: Ye, Zhijing, et al.
Published: (2025)
by: Ye, Zhijing, et al.
Published: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
by: Dai, Song, et al.
Published: (2025)
by: Dai, Song, et al.
Published: (2025)
FEANEL: A Benchmark for Fine-Grained Error Analysis in K-12 English Writing
by: Ye, Jingheng, et al.
Published: (2025)
by: Ye, Jingheng, et al.
Published: (2025)
A Compact Model for English Grammar Error Correction in the Low‐Latency Edge Deployment
by: Shaoli Xiong
Published: (2026)
by: Shaoli Xiong
Published: (2026)
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
by: Zhou, Guanyu, et al.
Published: (2024)
by: Zhou, Guanyu, et al.
Published: (2024)
OneForecast: A Universal Framework for Global and Regional Weather Forecasting
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
Diff-GNSS: Diffusion-based Pseudorange Error Estimation
by: Zhu, Jiaqi, et al.
Published: (2025)
by: Zhu, Jiaqi, et al.
Published: (2025)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
by: Zheng, Kening, et al.
Published: (2024)
by: Zheng, Kening, et al.
Published: (2024)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
by: Wang, Shaoguang, et al.
Published: (2026)
by: Wang, Shaoguang, et al.
Published: (2026)
Chapter Galileo’s Mathematical Errors
by: Blåsjö, Viktor
Published: (2024)
by: Blåsjö, Viktor
Published: (2024)
UniEDU: A Unified Language and Vision Assistant for Education Applications
by: Chu, Zhendong, et al.
Published: (2025)
by: Chu, Zhendong, et al.
Published: (2025)
Error‐Corrected Eternal Lifetime Storage
by: Jie Ma, et al.
Published: (2025)
by: Jie Ma, et al.
Published: (2025)
Error-Corrected Eternal Lifetime Storage
by: Ma, Jie, et al.
Published: (2025)
by: Ma, Jie, et al.
Published: (2025)
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
by: Xu, Zhaopan, et al.
Published: (2025)
by: Xu, Zhaopan, et al.
Published: (2025)
Similar Items
-
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection
by: Yan, Yibo, et al.
Published: (2025) -
Multimodal AI Teacher: Integrating Edge Computing and Reasoning Models for Enhanced Student Error Analysis
by: Tianlong Xu, et al.
Published: (2025) -
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
by: Yan, Yibo, et al.
Published: (2024) -
AI-Driven Virtual Teacher for Enhanced Educational Efficiency: Leveraging Large Pretrain Models for Autonomous Error Analysis and Correction
by: Xu, Tianlong, et al.
Published: (2024) -
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
by: Yan, Yibo, et al.
Published: (2025)