Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sáez, Arnau Igualde, Rhomrasi, Lamyae, Ahsini, Yusef, Vinuesa, Ricardo, Hoyas, Sergio, Sabater, Jose P. García, Alfonso, Marius J. Fullana i, Conejero, J. Alberto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation
von: Alanazi, Ahmed, et al.
Veröffentlicht: (2025)
von: Alanazi, Ahmed, et al.
Veröffentlicht: (2025)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
von: Chen, Kewei, et al.
Veröffentlicht: (2025)
von: Chen, Kewei, et al.
Veröffentlicht: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
von: Chen, Kewei, et al.
Veröffentlicht: (2026)
von: Chen, Kewei, et al.
Veröffentlicht: (2026)
Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning
von: Koubaa, Anis, et al.
Veröffentlicht: (2025)
von: Koubaa, Anis, et al.
Veröffentlicht: (2025)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
von: Zhang, Junwen, et al.
Veröffentlicht: (2025)
von: Zhang, Junwen, et al.
Veröffentlicht: (2025)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
von: Sharma, Aditya, et al.
Veröffentlicht: (2025)
von: Sharma, Aditya, et al.
Veröffentlicht: (2025)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
von: Chahine, Makram, et al.
Veröffentlicht: (2024)
von: Chahine, Makram, et al.
Veröffentlicht: (2024)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
von: Koh, Hyunseo, et al.
Veröffentlicht: (2026)
von: Koh, Hyunseo, et al.
Veröffentlicht: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
Tricks and Plug-ins for Gradient Boosting in Image Classification
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous Driving
von: Da, Longchao, et al.
Veröffentlicht: (2025)
von: Da, Longchao, et al.
Veröffentlicht: (2025)
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
von: Zeng, Zhitao, et al.
Veröffentlicht: (2026)
von: Zeng, Zhitao, et al.
Veröffentlicht: (2026)
MaP-AVR: A Meta-Action Planner for Agents Leveraging Vision Language Models and Retrieval-Augmented Generation
von: Guo, Zhenglong, et al.
Veröffentlicht: (2025)
von: Guo, Zhenglong, et al.
Veröffentlicht: (2025)
VGA: Vision GUI Assistant -- Minimizing Hallucinations through Image-Centric Fine-Tuning
von: Meng, Ziyang, et al.
Veröffentlicht: (2024)
von: Meng, Ziyang, et al.
Veröffentlicht: (2024)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
von: Qi, Dekang, et al.
Veröffentlicht: (2026)
von: Qi, Dekang, et al.
Veröffentlicht: (2026)
A Landmark-Aware Visual Navigation Dataset
von: Johnson, Faith, et al.
Veröffentlicht: (2024)
von: Johnson, Faith, et al.
Veröffentlicht: (2024)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
von: Menon, Anjali R., et al.
Veröffentlicht: (2025)
von: Menon, Anjali R., et al.
Veröffentlicht: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
von: Zhao, Lepeng, et al.
Veröffentlicht: (2026)
von: Zhao, Lepeng, et al.
Veröffentlicht: (2026)
Multi-Scale Graph Learning for Anti-Sparse Downscaling
von: Fan, Yingda, et al.
Veröffentlicht: (2025)
von: Fan, Yingda, et al.
Veröffentlicht: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
von: Adžemović, Momir
Veröffentlicht: (2025)
von: Adžemović, Momir
Veröffentlicht: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
Order-Robust Class Incremental Learning: Graph-Driven Dynamic Similarity Grouping
von: Lai, Guannan, et al.
Veröffentlicht: (2025)
von: Lai, Guannan, et al.
Veröffentlicht: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
von: Farzulla, Murad
Veröffentlicht: (2026)
von: Farzulla, Murad
Veröffentlicht: (2026)
Key-Scan-Based Mobile Robot Navigation: Integrated Mapping, Planning, and Control using Graphs of Scan Regions
von: Latha, Dharshan Bashkaran, et al.
Veröffentlicht: (2024)
von: Latha, Dharshan Bashkaran, et al.
Veröffentlicht: (2024)
When Does Global Attention Help? A Unified Empirical Study on Atomistic Graph Learning
von: Chowdhury, Arindam, et al.
Veröffentlicht: (2025)
von: Chowdhury, Arindam, et al.
Veröffentlicht: (2025)
RoboMoRe: LLM-based Robot Co-design via Joint Optimization of Morphology and Reward
von: Fang, Jiawei, et al.
Veröffentlicht: (2025)
von: Fang, Jiawei, et al.
Veröffentlicht: (2025)
Method of UAV Inspection of Photovoltaic Modules Using Thermal and RGB Data Fusion
von: Lysyi, Andrii, et al.
Veröffentlicht: (2025)
von: Lysyi, Andrii, et al.
Veröffentlicht: (2025)
Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
von: Wang, Wentao, et al.
Veröffentlicht: (2025)
von: Wang, Wentao, et al.
Veröffentlicht: (2025)
Banana Ripeness Level Classification using a Simple CNN Model Trained with Real and Synthetic Datasets
von: Chuquimarca, Luis, et al.
Veröffentlicht: (2025)
von: Chuquimarca, Luis, et al.
Veröffentlicht: (2025)
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
von: Koukounas, Andreas, et al.
Veröffentlicht: (2024)
von: Koukounas, Andreas, et al.
Veröffentlicht: (2024)
GLL: A Differentiable Graph Learning Layer for Neural Networks
von: Brown, Jason, et al.
Veröffentlicht: (2024)
von: Brown, Jason, et al.
Veröffentlicht: (2024)
CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models
von: Yao, Dongyu, et al.
Veröffentlicht: (2024)
von: Yao, Dongyu, et al.
Veröffentlicht: (2024)
RACAS: Controlling Diverse Robots With a Single Agentic System
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
von: Marmoret, Axel, et al.
Veröffentlicht: (2025)
von: Marmoret, Axel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
von: Viveiros, André G., et al.
Veröffentlicht: (2025) -
PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation
von: Alanazi, Ahmed, et al.
Veröffentlicht: (2025) -
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
von: Chen, Kewei, et al.
Veröffentlicht: (2025) -
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
von: Chen, Kewei, et al.
Veröffentlicht: (2026) -
Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning
von: Koubaa, Anis, et al.
Veröffentlicht: (2025)