Saved in:
Bibliographic Details
Main Authors: Zhou, Yitong, Cheng, Mingyue, Mao, Qingyang, Xu, Feiyang, Li, Xin
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2412.20662
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915312820027392
author Zhou, Yitong
Cheng, Mingyue
Mao, Qingyang
Xu, Feiyang
Li, Xin
author_facet Zhou, Yitong
Cheng, Mingyue
Mao, Qingyang
Xu, Feiyang
Li, Xin
contents Pre-trained foundation models have recently made significant progress in table-related tasks such as table understanding and reasoning. However, recognizing the structure and content of unstructured tables using Vision Large Language Models (VLLMs) remains under-explored. To bridge this gap, we propose a benchmark based on a hierarchical design philosophy to evaluate the recognition capabilities of VLLMs in training-free scenarios. Through in-depth evaluations, we find that low-quality image input is a significant bottleneck in the recognition process. Drawing inspiration from this, we propose the Neighbor-Guided Toolchain Reasoner (NGTR) framework, which is characterized by integrating diverse lightweight tools for visual operations aimed at mitigating issues with low-quality images. Specifically, we transfer a tool selection experience from a similar neighbor to the input and design a reflection module to supervise the tool invocation process. Extensive experiments on public datasets demonstrate that our approach significantly enhances the recognition capabilities of the vanilla VLLMs. We believe that the benchmark and framework could provide an alternative solution to table recognition.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20662
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Table Recognition with Vision LLMs: A Benchmark and Neighbor-Guided Toolchain Reasoner
Zhou, Yitong
Cheng, Mingyue
Mao, Qingyang
Xu, Feiyang
Li, Xin
Computer Vision and Pattern Recognition
Artificial Intelligence
Pre-trained foundation models have recently made significant progress in table-related tasks such as table understanding and reasoning. However, recognizing the structure and content of unstructured tables using Vision Large Language Models (VLLMs) remains under-explored. To bridge this gap, we propose a benchmark based on a hierarchical design philosophy to evaluate the recognition capabilities of VLLMs in training-free scenarios. Through in-depth evaluations, we find that low-quality image input is a significant bottleneck in the recognition process. Drawing inspiration from this, we propose the Neighbor-Guided Toolchain Reasoner (NGTR) framework, which is characterized by integrating diverse lightweight tools for visual operations aimed at mitigating issues with low-quality images. Specifically, we transfer a tool selection experience from a similar neighbor to the input and design a reflection module to supervise the tool invocation process. Extensive experiments on public datasets demonstrate that our approach significantly enhances the recognition capabilities of the vanilla VLLMs. We believe that the benchmark and framework could provide an alternative solution to table recognition.
title Enhancing Table Recognition with Vision LLMs: A Benchmark and Neighbor-Guided Toolchain Reasoner
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.20662