Every Part Matters: Integrity Verification of Scientific Figures Based on Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Xiang, Liu, Jiawei, Liu, Yinpeng, Cheng, Qikai, Lu, Wei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914888503263232
author Shi, Xiang
Liu, Jiawei
Liu, Yinpeng
Cheng, Qikai
Lu, Wei
author_facet Shi, Xiang
Liu, Jiawei
Liu, Yinpeng
Cheng, Qikai
Lu, Wei
contents This paper tackles a key issue in the interpretation of scientific figures: the fine-grained alignment of text and figures. It advances beyond prior research that primarily dealt with straightforward, data-driven visualizations such as bar and pie charts and only offered a basic understanding of diagrams through captioning and classification. We introduce a novel task, Figure Integrity Verification, designed to evaluate the precision of technologies in aligning textual knowledge with visual elements in scientific figures. To support this, we develop a semi-automated method for constructing a large-scale dataset, Figure-seg, specifically designed for this task. Additionally, we propose an innovative framework, Every Part Matters (EPM), which leverages Multimodal Large Language Models (MLLMs) to not only incrementally improve the alignment and verification of text-figure integrity but also enhance integrity through analogical reasoning. Our comprehensive experiments show that these innovations substantially improve upon existing methods, allowing for more precise and thorough analysis of complex scientific figures. This progress not only enhances our understanding of multimodal technologies but also stimulates further research and practical applications across fields requiring the accurate interpretation of complex visual data.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18626
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Every Part Matters: Integrity Verification of Scientific Figures Based on Multimodal Large Language Models
Shi, Xiang
Liu, Jiawei
Liu, Yinpeng
Cheng, Qikai
Lu, Wei
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Digital Libraries
Multimedia
This paper tackles a key issue in the interpretation of scientific figures: the fine-grained alignment of text and figures. It advances beyond prior research that primarily dealt with straightforward, data-driven visualizations such as bar and pie charts and only offered a basic understanding of diagrams through captioning and classification. We introduce a novel task, Figure Integrity Verification, designed to evaluate the precision of technologies in aligning textual knowledge with visual elements in scientific figures. To support this, we develop a semi-automated method for constructing a large-scale dataset, Figure-seg, specifically designed for this task. Additionally, we propose an innovative framework, Every Part Matters (EPM), which leverages Multimodal Large Language Models (MLLMs) to not only incrementally improve the alignment and verification of text-figure integrity but also enhance integrity through analogical reasoning. Our comprehensive experiments show that these innovations substantially improve upon existing methods, allowing for more precise and thorough analysis of complex scientific figures. This progress not only enhances our understanding of multimodal technologies but also stimulates further research and practical applications across fields requiring the accurate interpretation of complex visual data.
title Every Part Matters: Integrity Verification of Scientific Figures Based on Multimodal Large Language Models
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Digital Libraries
Multimedia
url https://arxiv.org/abs/2407.18626