Can Vision-Language Models Evaluate Handwritten Math?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nath, Oikantik, Bathina, Hanani, Khan, Mohammed Safi Ur Rahman, Khapra, Mitesh M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2026)
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2026)
IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
von: Doddapaneni, Sumanth, et al.
Veröffentlicht: (2024)
von: Doddapaneni, Sumanth, et al.
Veröffentlicht: (2024)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
Automated LaTeX Code Generation from Handwritten Math Expressions Using Vision Transformer
von: Sundararaj, Jayaprakash, et al.
Veröffentlicht: (2024)
von: Sundararaj, Jayaprakash, et al.
Veröffentlicht: (2024)
Can Vision-Language Models Solve Visual Math Equations?
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
von: Baral, Sami, et al.
Veröffentlicht: (2025)
von: Baral, Sami, et al.
Veröffentlicht: (2025)
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
von: Doddapaneni, Sumanth, et al.
Veröffentlicht: (2024)
von: Doddapaneni, Sumanth, et al.
Veröffentlicht: (2024)
FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes
von: Nawale, Janki Atul, et al.
Veröffentlicht: (2025)
von: Nawale, Janki Atul, et al.
Veröffentlicht: (2025)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
von: Yuan, Fan, et al.
Veröffentlicht: (2025)
von: Yuan, Fan, et al.
Veröffentlicht: (2025)
Can Vision-Language Models Solve the Shell Game?
von: Liu, Tiedong, et al.
Veröffentlicht: (2026)
von: Liu, Tiedong, et al.
Veröffentlicht: (2026)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
von: Lucy, Li, et al.
Veröffentlicht: (2026)
von: Lucy, Li, et al.
Veröffentlicht: (2026)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
Evaluating Vision-Language Models as Evaluators in Path Planning
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2024)
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2024)
Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
von: Song, Dingjie, et al.
Veröffentlicht: (2026)
von: Song, Dingjie, et al.
Veröffentlicht: (2026)
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
von: Gafni, Tomer, et al.
Veröffentlicht: (2025)
von: Gafni, Tomer, et al.
Veröffentlicht: (2025)
Evaluating Vision-Language Models for Emotion Recognition
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2025)
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2025)
Evaluation of Cultural Competence of Vision-Language Models
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
von: Nath, Sujoy, et al.
Veröffentlicht: (2025)
von: Nath, Sujoy, et al.
Veröffentlicht: (2025)
Seeing Justice Clearly: Handwritten Legal Document Translation with OCR and Vision-Language Models
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
Can We Predict Performance of Large Models across Vision-Language Tasks?
von: Zhao, Qinyu, et al.
Veröffentlicht: (2024)
von: Zhao, Qinyu, et al.
Veröffentlicht: (2024)
BHDD: A Burmese Handwritten Digit Dataset
von: Aung, Swan Htet, et al.
Veröffentlicht: (2026)
von: Aung, Swan Htet, et al.
Veröffentlicht: (2026)
Instruction-Following Evaluation of Large Vision-Language Models
von: Shiono, Daiki, et al.
Veröffentlicht: (2025)
von: Shiono, Daiki, et al.
Veröffentlicht: (2025)
ColorBlindnessEval: Can Vision-Language Models Pass Color Blindness Tests?
von: Ling, Zijian, et al.
Veröffentlicht: (2025)
von: Ling, Zijian, et al.
Veröffentlicht: (2025)
Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
von: Sun, Min Woo, et al.
Veröffentlicht: (2025)
von: Sun, Min Woo, et al.
Veröffentlicht: (2025)
Revisiting N-Gram Models: Their Impact in Modern Neural Networks for Handwritten Text Recognition
von: Tarride, Solène, et al.
Veröffentlicht: (2024)
von: Tarride, Solène, et al.
Veröffentlicht: (2024)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
von: Lu, Jiaying, et al.
Veröffentlicht: (2023)
von: Lu, Jiaying, et al.
Veröffentlicht: (2023)
Can Vision Language Models Understand Mimed Actions?
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Evaluating Vision Language Model Adaptations for Radiology Report Generation in Low-Resource Languages
von: Salmè, Marco, et al.
Veröffentlicht: (2025)
von: Salmè, Marco, et al.
Veröffentlicht: (2025)
Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
HazardNet: A Small-Scale Vision Language Model for Real-Time Traffic Safety Detection at Edge Devices
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025)
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025)
The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form-Meaning Mapping
von: Keleş, Onur, et al.
Veröffentlicht: (2025)
von: Keleş, Onur, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2026) -
IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
von: Nath, Oikantik, et al.
Veröffentlicht: (2025) -
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
von: Doddapaneni, Sumanth, et al.
Veröffentlicht: (2024) -
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025) -
Automated LaTeX Code Generation from Handwritten Math Expressions Using Vision Transformer
von: Sundararaj, Jayaprakash, et al.
Veröffentlicht: (2024)