Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Yingjie, Bai, Xuefeng, Chen, Kehai, Xiang, Yang, Yu, Jun, Zhang, Min
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910990901182464
author Zhu, Yingjie
Bai, Xuefeng
Chen, Kehai
Xiang, Yang
Yu, Jun
Zhang, Min
author_facet Zhu, Yingjie
Bai, Xuefeng
Chen, Kehai
Xiang, Yang
Yu, Jun
Zhang, Min
contents Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across diverse tasks. Despite great success, recent studies show that LVLMs encounter substantial limitations when engaging with visual graphs. To study the reason behind these limitations, we propose VGCure, a comprehensive benchmark covering 22 tasks for examining the fundamental graph understanding and reasoning capacities of LVLMs. Extensive evaluations conducted on 14 LVLMs reveal that LVLMs are weak in basic graph understanding and reasoning tasks, particularly those concerning relational or structurally complex information. Based on this observation, we propose a structure-aware fine-tuning framework to enhance LVLMs with structure learning abilities through three self-supervised learning tasks. Experiments validate the effectiveness of our method in improving LVLMs' performance on fundamental and downstream graph learning tasks, as well as enhancing their robustness against complex visual graphs.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13540
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
Zhu, Yingjie
Bai, Xuefeng
Chen, Kehai
Xiang, Yang
Yu, Jun
Zhang, Min
Computation and Language
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across diverse tasks. Despite great success, recent studies show that LVLMs encounter substantial limitations when engaging with visual graphs. To study the reason behind these limitations, we propose VGCure, a comprehensive benchmark covering 22 tasks for examining the fundamental graph understanding and reasoning capacities of LVLMs. Extensive evaluations conducted on 14 LVLMs reveal that LVLMs are weak in basic graph understanding and reasoning tasks, particularly those concerning relational or structurally complex information. Based on this observation, we propose a structure-aware fine-tuning framework to enhance LVLMs with structure learning abilities through three self-supervised learning tasks. Experiments validate the effectiveness of our method in improving LVLMs' performance on fundamental and downstream graph learning tasks, as well as enhancing their robustness against complex visual graphs.
title Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.13540