DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baral, Sami, Lucy, Li, Knight, Ryan, Ng, Alice, Soldaini, Luca, Heffernan, Neil T., Lo, Kyle |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
von: Lucy, Li, et al.
Veröffentlicht: (2026)
von: Lucy, Li, et al.
Veröffentlicht: (2026)
Mathfish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula
von: Lucy, Li, et al.
Veröffentlicht: (2024)
von: Lucy, Li, et al.
Veröffentlicht: (2024)
Automated Feedback in Math Education: A Comparative Analysis of LLMs for Open-Ended Responses
von: Baral, Sami, et al.
Veröffentlicht: (2024)
von: Baral, Sami, et al.
Veröffentlicht: (2024)
olmOCR 2: Unit Test Rewards for Document OCR
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
Can Vision-Language Models Evaluate Handwritten Math?
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
A Multi-Agent Approach to Validate and Refine LLM-Generated Personalized Math Problems
von: Ikram, Fareya, et al.
Veröffentlicht: (2026)
von: Ikram, Fareya, et al.
Veröffentlicht: (2026)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
Using Java Geometry Expert as Guide in the Preparations for Math Contests
von: Ganglmayr, Ines, et al.
Veröffentlicht: (2024)
von: Ganglmayr, Ines, et al.
Veröffentlicht: (2024)
SafeMath: Inference-time Safety improves Math Accuracy
von: Basu, Sagnik, et al.
Veröffentlicht: (2026)
von: Basu, Sagnik, et al.
Veröffentlicht: (2026)
Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
von: Anantheswaran, Ujjwala, et al.
Veröffentlicht: (2024)
von: Anantheswaran, Ujjwala, et al.
Veröffentlicht: (2024)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
MathBuddy: A Multimodal System for Affective Math Tutoring
von: Kar, Debanjana, et al.
Veröffentlicht: (2025)
von: Kar, Debanjana, et al.
Veröffentlicht: (2025)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
CoinMath: Harnessing the Power of Coding Instruction for Math LLMs
von: Wei, Chengwei, et al.
Veröffentlicht: (2024)
von: Wei, Chengwei, et al.
Veröffentlicht: (2024)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
Automatic Detection of Research Values from Scientific Abstracts Across Computer Science Subfields
von: Jiang, Hang, et al.
Veröffentlicht: (2025)
von: Jiang, Hang, et al.
Veröffentlicht: (2025)
RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)
von: Cuclea, Luca-Ncolae, et al.
Veröffentlicht: (2026)
von: Cuclea, Luca-Ncolae, et al.
Veröffentlicht: (2026)
MATHWELL: Generating Educational Math Word Problems Using Teacher Annotations
von: Christ, Bryan R, et al.
Veröffentlicht: (2024)
von: Christ, Bryan R, et al.
Veröffentlicht: (2024)
Author Intent: Eliminating Ambiguity in MathML
von: Carlisle, David, et al.
Veröffentlicht: (2024)
von: Carlisle, David, et al.
Veröffentlicht: (2024)
The CompMath-MCQ Dataset: Are LLMs Ready for Higher-Level Math?
von: Raimondi, Bianca, et al.
Veröffentlicht: (2026)
von: Raimondi, Bianca, et al.
Veröffentlicht: (2026)
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
von: Liu, Wentao, et al.
Veröffentlicht: (2024)
von: Liu, Wentao, et al.
Veröffentlicht: (2024)
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
von: Zhang, Boning, et al.
Veröffentlicht: (2024)
von: Zhang, Boning, et al.
Veröffentlicht: (2024)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
von: Mitra, Arindam, et al.
Veröffentlicht: (2024)
von: Mitra, Arindam, et al.
Veröffentlicht: (2024)
MegaMath: Pushing the Limits of Open Math Corpora
von: Zhou, Fan, et al.
Veröffentlicht: (2025)
von: Zhou, Fan, et al.
Veröffentlicht: (2025)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
von: Van Long, Phuoc Pham, et al.
Veröffentlicht: (2023)
von: Van Long, Phuoc Pham, et al.
Veröffentlicht: (2023)
Can Vision-Language Models Solve Visual Math Equations?
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
Organize the Web: Constructing Domains Enhances Pre-Training Data Curation
von: Wettig, Alexander, et al.
Veröffentlicht: (2025)
von: Wettig, Alexander, et al.
Veröffentlicht: (2025)
KIWI: A Dataset of Knowledge-Intensive Writing Instructions for Answering Research Questions
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
MisEdu-RAG: A Misconception-Aware Dual-Hypergraph RAG for Novice Math Teachers
von: Guo, Zhihan, et al.
Veröffentlicht: (2026)
von: Guo, Zhihan, et al.
Veröffentlicht: (2026)
MathChat: Converse to Tackle Challenging Math Problems with LLM Agents
von: Wu, Yiran, et al.
Veröffentlicht: (2023)
von: Wu, Yiran, et al.
Veröffentlicht: (2023)
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams
von: Caraeni, Adriana, et al.
Veröffentlicht: (2024)
von: Caraeni, Adriana, et al.
Veröffentlicht: (2024)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
von: Ying, Huaiyuan, et al.
Veröffentlicht: (2024)
von: Ying, Huaiyuan, et al.
Veröffentlicht: (2024)
GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving
von: Zhang, Yingji, et al.
Veröffentlicht: (2026)
von: Zhang, Yingji, et al.
Veröffentlicht: (2026)
Tangible Math
von: Scarlatos, Lori L.
Veröffentlicht: (2006)
von: Scarlatos, Lori L.
Veröffentlicht: (2006)
Models Can and Should Embrace the Communicative Nature of Human-Generated Math
von: Boguraev, Sasha, et al.
Veröffentlicht: (2024)
von: Boguraev, Sasha, et al.
Veröffentlicht: (2024)
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
von: Tian, Shi-Yu, et al.
Veröffentlicht: (2025)
von: Tian, Shi-Yu, et al.
Veröffentlicht: (2025)
Effective and Scalable Math Support: Evidence on the Impact of an AI- Tutor on Math Achievement in Ghana
von: Henkel, Owen, et al.
Veröffentlicht: (2024)
von: Henkel, Owen, et al.
Veröffentlicht: (2024)
MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
von: Li, Chengpeng, et al.
Veröffentlicht: (2023)
von: Li, Chengpeng, et al.
Veröffentlicht: (2023)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
von: Wang, Zengzhi, et al.
Veröffentlicht: (2023)
von: Wang, Zengzhi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
von: Lucy, Li, et al.
Veröffentlicht: (2026) -
Mathfish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula
von: Lucy, Li, et al.
Veröffentlicht: (2024) -
Automated Feedback in Math Education: A Comparative Analysis of LLMs for Open-Ended Responses
von: Baral, Sami, et al.
Veröffentlicht: (2024) -
olmOCR 2: Unit Test Rewards for Document OCR
von: Poznanski, Jake, et al.
Veröffentlicht: (2025) -
Can Vision-Language Models Evaluate Handwritten Math?
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)