NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xiang, Hua, Wenyue, Zhu, Kaijie, Li, Lingyao, Ling, Haoyang, Chi, Jinkui, Dou, Qi, Wang, Jindong, Zhang, Yongfeng, Ma, Xin, Fan, Lizhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915464529051648
author Li, Xiang
Hua, Wenyue
Zhu, Kaijie
Li, Lingyao
Ling, Haoyang
Chi, Jinkui
Dou, Qi
Wang, Jindong
Zhang, Yongfeng
Ma, Xin
Fan, Lizhou
author_facet Li, Xiang
Hua, Wenyue
Zhu, Kaijie
Li, Lingyao
Ling, Haoyang
Chi, Jinkui
Dou, Qi
Wang, Jindong
Zhang, Yongfeng
Ma, Xin
Fan, Lizhou
contents Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multimodal understanding, yet their reasoning abilities remain underexplored. Existing benchmarks tend to focus on perception or text-based comprehension, offering limited insight into how well these models perform on structured, logic-driven tasks that require both visual and linguistic reasoning. To address this gap, we introduce NPHardEval4V, a multimodal benchmark suite grounded in four classical NP-hard problems: Knapsack, Set Cover, Traveling Salesperson, and Vertex Cover. Each task is presented through a combination of structured visual layouts and textual prompts, designed to assess the ability of LVLMs to perform combinatorial reasoning under visual-linguistic constraints. We evaluate a set of advanced open-source and closed-source vision-language models under a unified prompting and problem representation framework. This enables fair comparison across models and task types, while isolating key variables affecting performance. Our results show that while these models perform reasonably well on perception-based inputs, they struggle with global optimization, abstraction, and constraint satisfaction. No single model demonstrates consistent reasoning capability across all problem types, and common failure patterns reveal fundamental limitations in current architectures. By leveraging the structure and complexity of NP-hard problems, NPHardEval4V provides a scalable, interpretable, and challenging testbed for diagnosing reasoning behaviors in LVLMs. We hope this benchmark can support the community in building more robust, inference-capable multimodal systems. The benchmark dataset and code are available at https://github.com/lizhouf/NPHardEval4.
format Preprint
id arxiv_https___arxiv_org_abs_2403_01777
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
Li, Xiang
Hua, Wenyue
Zhu, Kaijie
Li, Lingyao
Ling, Haoyang
Chi, Jinkui
Dou, Qi
Wang, Jindong
Zhang, Yongfeng
Ma, Xin
Fan, Lizhou
Computation and Language
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multimodal understanding, yet their reasoning abilities remain underexplored. Existing benchmarks tend to focus on perception or text-based comprehension, offering limited insight into how well these models perform on structured, logic-driven tasks that require both visual and linguistic reasoning. To address this gap, we introduce NPHardEval4V, a multimodal benchmark suite grounded in four classical NP-hard problems: Knapsack, Set Cover, Traveling Salesperson, and Vertex Cover. Each task is presented through a combination of structured visual layouts and textual prompts, designed to assess the ability of LVLMs to perform combinatorial reasoning under visual-linguistic constraints. We evaluate a set of advanced open-source and closed-source vision-language models under a unified prompting and problem representation framework. This enables fair comparison across models and task types, while isolating key variables affecting performance. Our results show that while these models perform reasonably well on perception-based inputs, they struggle with global optimization, abstraction, and constraint satisfaction. No single model demonstrates consistent reasoning capability across all problem types, and common failure patterns reveal fundamental limitations in current architectures. By leveraging the structure and complexity of NP-hard problems, NPHardEval4V provides a scalable, interpretable, and challenging testbed for diagnosing reasoning behaviors in LVLMs. We hope this benchmark can support the community in building more robust, inference-capable multimodal systems. The benchmark dataset and code are available at https://github.com/lizhouf/NPHardEval4.
title NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.01777