Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sáez, Arnau Igualde, Rhomrasi, Lamyae, Ahsini, Yusef, Vinuesa, Ricardo, Hoyas, Sergio, Sabater, Jose P. García, Alfonso, Marius J. Fullana i, Conejero, J. Alberto
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910995530645504
author Sáez, Arnau Igualde
Rhomrasi, Lamyae
Ahsini, Yusef
Vinuesa, Ricardo
Hoyas, Sergio
Sabater, Jose P. García
Alfonso, Marius J. Fullana i
Conejero, J. Alberto
author_facet Sáez, Arnau Igualde
Rhomrasi, Lamyae
Ahsini, Yusef
Vinuesa, Ricardo
Hoyas, Sergio
Sabater, Jose P. García
Alfonso, Marius J. Fullana i
Conejero, J. Alberto
contents Multimodal Large Language Models (MLLMs) promise advanced vision language capabilities, yet their effectiveness in visually presented mathematics remains underexplored. This paper analyzes the development and evaluation of MLLMs for mathematical problem solving, focusing on diagrams, multilingual text, and symbolic notation. We then assess several models, including GPT 4o, Pixtral, Qwen VL, Llama 3.2 Vision variants, and Gemini 2.0 Flash in a multilingual Kangaroo style benchmark spanning English, French, Spanish, and Catalan. Our experiments reveal four key findings. First, overall precision remains moderate across geometry, visual algebra, logic, patterns, and combinatorics: no single model excels in every topic. Second, while most models see improved accuracy with questions that do not have images, the gain is often limited; performance for some remains nearly unchanged without visual input, indicating underutilization of diagrammatic information. Third, substantial variation exists across languages and difficulty levels: models frequently handle easier items but struggle with advanced geometry and combinatorial reasoning. Notably, Gemini 2.0 Flash achieves the highest precision on image based tasks, followed by Qwen VL 2.5 72B and GPT 4o, though none approach human level performance. Fourth, a complementary analysis aimed at distinguishing whether models reason or simply recite reveals that Gemini and GPT 4o stand out for their structured reasoning and consistent accuracy. In contrast, Pixtral and Llama exhibit less consistent reasoning, often defaulting to heuristics or randomness when unable to align their outputs with the given answer options.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07418
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
Sáez, Arnau Igualde
Rhomrasi, Lamyae
Ahsini, Yusef
Vinuesa, Ricardo
Hoyas, Sergio
Sabater, Jose P. García
Alfonso, Marius J. Fullana i
Conejero, J. Alberto
Artificial Intelligence
68T05, 68T45
I.2.10; I.2.7
Multimodal Large Language Models (MLLMs) promise advanced vision language capabilities, yet their effectiveness in visually presented mathematics remains underexplored. This paper analyzes the development and evaluation of MLLMs for mathematical problem solving, focusing on diagrams, multilingual text, and symbolic notation. We then assess several models, including GPT 4o, Pixtral, Qwen VL, Llama 3.2 Vision variants, and Gemini 2.0 Flash in a multilingual Kangaroo style benchmark spanning English, French, Spanish, and Catalan. Our experiments reveal four key findings. First, overall precision remains moderate across geometry, visual algebra, logic, patterns, and combinatorics: no single model excels in every topic. Second, while most models see improved accuracy with questions that do not have images, the gain is often limited; performance for some remains nearly unchanged without visual input, indicating underutilization of diagrammatic information. Third, substantial variation exists across languages and difficulty levels: models frequently handle easier items but struggle with advanced geometry and combinatorial reasoning. Notably, Gemini 2.0 Flash achieves the highest precision on image based tasks, followed by Qwen VL 2.5 72B and GPT 4o, though none approach human level performance. Fourth, a complementary analysis aimed at distinguishing whether models reason or simply recite reveals that Gemini and GPT 4o stand out for their structured reasoning and consistent accuracy. In contrast, Pixtral and Llama exhibit less consistent reasoning, often defaulting to heuristics or randomness when unable to align their outputs with the given answer options.
title Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
topic Artificial Intelligence
68T05, 68T45
I.2.10; I.2.7
url https://arxiv.org/abs/2506.07418