Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
Fuente:
arXiv
Salvato in:
| Autori principali: | Cherian, Anoop, Peng, Kuan-Chuan, Lohit, Suhas, Matthiesen, Joanna, Smith, Kevin, Tenenbaum, Joshua B. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
di: Cherian, Anoop, et al.
Pubblicazione: (2025)
di: Cherian, Anoop, et al.
Pubblicazione: (2025)
Auto-Vocabulary 3D Object Detection
di: Zhang, Haomeng, et al.
Pubblicazione: (2025)
di: Zhang, Haomeng, et al.
Pubblicazione: (2025)
Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
di: Cui, Yiming, et al.
Pubblicazione: (2025)
di: Cui, Yiming, et al.
Pubblicazione: (2025)
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
di: Xiang, Xinhao, et al.
Pubblicazione: (2025)
di: Xiang, Xinhao, et al.
Pubblicazione: (2025)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
di: Chen, Qiguang, et al.
Pubblicazione: (2026)
di: Chen, Qiguang, et al.
Pubblicazione: (2026)
Equivariant Spatio-Temporal Self-Supervision for LiDAR Object Detection
di: Hegde, Deepti, et al.
Pubblicazione: (2024)
di: Hegde, Deepti, et al.
Pubblicazione: (2024)
Multimodal 3D Object Detection on Unseen Domains
di: Hegde, Deepti, et al.
Pubblicazione: (2024)
di: Hegde, Deepti, et al.
Pubblicazione: (2024)
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
di: Li, Danrui, et al.
Pubblicazione: (2026)
di: Li, Danrui, et al.
Pubblicazione: (2026)
TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models
di: Ni, Haomiao, et al.
Pubblicazione: (2024)
di: Ni, Haomiao, et al.
Pubblicazione: (2024)
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
di: Qiu, Haoyi, et al.
Pubblicazione: (2024)
di: Qiu, Haoyi, et al.
Pubblicazione: (2024)
Programmatic Video Prediction Using Large Language Models
di: Tang, Hao, et al.
Pubblicazione: (2025)
di: Tang, Hao, et al.
Pubblicazione: (2025)
Improving Open-World Object Localization by Discovering Background
di: Singh, Ashish, et al.
Pubblicazione: (2025)
di: Singh, Ashish, et al.
Pubblicazione: (2025)
Instruction-Following Evaluation of Large Vision-Language Models
di: Shiono, Daiki, et al.
Pubblicazione: (2025)
di: Shiono, Daiki, et al.
Pubblicazione: (2025)
Building Cooperative Embodied Agents Modularly with Large Language Models
di: Zhang, Hongxin, et al.
Pubblicazione: (2023)
di: Zhang, Hongxin, et al.
Pubblicazione: (2023)
Malayalam Sign Language Identification using Finetuned YOLOv8 and Computer Vision Techniques
di: K., Abhinand, et al.
Pubblicazione: (2024)
di: K., Abhinand, et al.
Pubblicazione: (2024)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
di: Lu, Jiaying, et al.
Pubblicazione: (2023)
di: Lu, Jiaying, et al.
Pubblicazione: (2023)
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
di: Li, Yue, et al.
Pubblicazione: (2025)
di: Li, Yue, et al.
Pubblicazione: (2025)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
di: Wang, Zihu, et al.
Pubblicazione: (2025)
di: Wang, Zihu, et al.
Pubblicazione: (2025)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
di: Wu, Yuhang, et al.
Pubblicazione: (2024)
di: Wu, Yuhang, et al.
Pubblicazione: (2024)
Neuro-Symbolic Concepts
di: Mao, Jiayuan, et al.
Pubblicazione: (2025)
di: Mao, Jiayuan, et al.
Pubblicazione: (2025)
Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models
di: Lei, Xuanyu, et al.
Pubblicazione: (2024)
di: Lei, Xuanyu, et al.
Pubblicazione: (2024)
VLMInferSlow: Evaluating the Efficiency Robustness of Large Vision-Language Models as a Service
di: Wang, Xiasi, et al.
Pubblicazione: (2025)
di: Wang, Xiasi, et al.
Pubblicazione: (2025)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
di: Wu, Xuyang, et al.
Pubblicazione: (2024)
di: Wu, Xuyang, et al.
Pubblicazione: (2024)
STAR: A Benchmark for Situated Reasoning in Real-World Videos
di: Wu, Bo, et al.
Pubblicazione: (2024)
di: Wu, Bo, et al.
Pubblicazione: (2024)
Evaluating Vision-Language Models as Evaluators in Path Planning
di: Aghzal, Mohamed, et al.
Pubblicazione: (2024)
di: Aghzal, Mohamed, et al.
Pubblicazione: (2024)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
di: Gu, Jihao, et al.
Pubblicazione: (2025)
di: Gu, Jihao, et al.
Pubblicazione: (2025)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
di: Li, Baiqi, et al.
Pubblicazione: (2024)
di: Li, Baiqi, et al.
Pubblicazione: (2024)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
di: Zou, Chengke, et al.
Pubblicazione: (2024)
di: Zou, Chengke, et al.
Pubblicazione: (2024)
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
di: Zhang, Junyi, et al.
Pubblicazione: (2025)
di: Zhang, Junyi, et al.
Pubblicazione: (2025)
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
di: Cheng, Sijie, et al.
Pubblicazione: (2023)
di: Cheng, Sijie, et al.
Pubblicazione: (2023)
Evaluating Vision-Language Models for Emotion Recognition
di: Bhattacharyya, Sree, et al.
Pubblicazione: (2025)
di: Bhattacharyya, Sree, et al.
Pubblicazione: (2025)
Evaluation of Cultural Competence of Vision-Language Models
di: Yadav, Srishti, et al.
Pubblicazione: (2025)
di: Yadav, Srishti, et al.
Pubblicazione: (2025)
Intriguing Properties of Large Language and Vision Models
di: Lee, Young-Jun, et al.
Pubblicazione: (2024)
di: Lee, Young-Jun, et al.
Pubblicazione: (2024)
Multimodal Diffusion Bridge with Attention-Based SAR Fusion for Satellite Image Cloud Removal
di: Hu, Yuyang, et al.
Pubblicazione: (2025)
di: Hu, Yuyang, et al.
Pubblicazione: (2025)
Overconfidence is Key: Verbalized Uncertainty Evaluation in Large Language and Vision-Language Models
di: Groot, Tobias, et al.
Pubblicazione: (2024)
di: Groot, Tobias, et al.
Pubblicazione: (2024)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
di: Liang, Hao, et al.
Pubblicazione: (2024)
di: Liang, Hao, et al.
Pubblicazione: (2024)
Can Vision-Language Models Evaluate Handwritten Math?
di: Nath, Oikantik, et al.
Pubblicazione: (2025)
di: Nath, Oikantik, et al.
Pubblicazione: (2025)
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
di: Tang, Liyan, et al.
Pubblicazione: (2025)
di: Tang, Liyan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
di: Cherian, Anoop, et al.
Pubblicazione: (2025) -
Auto-Vocabulary 3D Object Detection
di: Zhang, Haomeng, et al.
Pubblicazione: (2025) -
Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
di: Cui, Yiming, et al.
Pubblicazione: (2025) -
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
di: Xiang, Xinhao, et al.
Pubblicazione: (2025) -
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
di: Chen, Qiguang, et al.
Pubblicazione: (2026)