Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Srivastava, Archita, Kumar, Abhas, Kumar, Rajesh, Srinivasan, Prabhakar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913641596452864
author Srivastava, Archita
Kumar, Abhas
Kumar, Rajesh
Srinivasan, Prabhakar
author_facet Srivastava, Archita
Kumar, Abhas
Kumar, Rajesh
Srinivasan, Prabhakar
contents Chart interpretation is crucial for visual data analysis, but accurately extracting information from charts poses significant challenges for automated models. This study investigates the fine-tuning of DEPLOT, a modality conversion module that translates the image of a plot or chart to a linearized table, on a custom dataset of 50,000 bar charts. The dataset comprises simple, stacked, and grouped bar charts, targeting the unique structural features of these visualizations. The finetuned DEPLOT model is evaluated against its base version using a test set of 1,000 images and two metrics: Relative Mapping Similarity (RMS), which measures categorical mapping accuracy, and Relative Number Set Similarity (RNSS), which evaluates numerical interpretation accuracy. To further explore the reasoning capabilities of large language models (LLMs), we curate an additional set of 100 bar chart images paired with question answer sets. Our findings demonstrate that providing a structured intermediate table alongside the image significantly enhances LLM reasoning performance compared to direct image queries.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04675
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
Srivastava, Archita
Kumar, Abhas
Kumar, Rajesh
Srinivasan, Prabhakar
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Chart interpretation is crucial for visual data analysis, but accurately extracting information from charts poses significant challenges for automated models. This study investigates the fine-tuning of DEPLOT, a modality conversion module that translates the image of a plot or chart to a linearized table, on a custom dataset of 50,000 bar charts. The dataset comprises simple, stacked, and grouped bar charts, targeting the unique structural features of these visualizations. The finetuned DEPLOT model is evaluated against its base version using a test set of 1,000 images and two metrics: Relative Mapping Similarity (RMS), which measures categorical mapping accuracy, and Relative Number Set Similarity (RNSS), which evaluates numerical interpretation accuracy. To further explore the reasoning capabilities of large language models (LLMs), we curate an additional set of 100 bar chart images paired with question answer sets. Our findings demonstrate that providing a structured intermediate table alongside the image significantly enhances LLM reasoning performance compared to direct image queries.
title Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2501.04675