Augmenting a Large Language Model with a Combination of Text and Visual Data for Conversational Visualization of Global Geospatial Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mena, Omar, Kouyoumdjian, Alexandre, Besançon, Lonni, Gleicher, Michael, Viola, Ivan, Ynnerman, Anders
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913653453750272
author Mena, Omar
Kouyoumdjian, Alexandre
Besançon, Lonni
Gleicher, Michael
Viola, Ivan
Ynnerman, Anders
author_facet Mena, Omar
Kouyoumdjian, Alexandre
Besançon, Lonni
Gleicher, Michael
Viola, Ivan
Ynnerman, Anders
contents We present a method for augmenting a Large Language Model (LLM) with a combination of text and visual data to enable accurate question answering in visualization of scientific data, making conversational visualization possible. LLMs struggle with tasks like visual data interaction, as they lack contextual visual information. We address this problem by merging a text description of a visualization and dataset with snapshots of the visualization. We extract their essential features into a structured text file, highly compact, yet descriptive enough to appropriately augment the LLM with contextual information, without any fine-tuning. This approach can be applied to any visualization that is already finally rendered, as long as it is associated with some textual description.
format Preprint
id arxiv_https___arxiv_org_abs_2501_09521
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Augmenting a Large Language Model with a Combination of Text and Visual Data for Conversational Visualization of Global Geospatial Data
Mena, Omar
Kouyoumdjian, Alexandre
Besançon, Lonni
Gleicher, Michael
Viola, Ivan
Ynnerman, Anders
Human-Computer Interaction
Computation and Language
We present a method for augmenting a Large Language Model (LLM) with a combination of text and visual data to enable accurate question answering in visualization of scientific data, making conversational visualization possible. LLMs struggle with tasks like visual data interaction, as they lack contextual visual information. We address this problem by merging a text description of a visualization and dataset with snapshots of the visualization. We extract their essential features into a structured text file, highly compact, yet descriptive enough to appropriately augment the LLM with contextual information, without any fine-tuning. This approach can be applied to any visualization that is already finally rendered, as long as it is associated with some textual description.
title Augmenting a Large Language Model with a Combination of Text and Visual Data for Conversational Visualization of Global Geospatial Data
topic Human-Computer Interaction
Computation and Language
url https://arxiv.org/abs/2501.09521