Generating Analytic Specifications for Data Visualization from Natural Language Queries using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sah, Subham, Mitra, Rishab, Narechania, Arpit, Endert, Alex, Stasko, John, Dou, Wenwen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913481627795456
author Sah, Subham
Mitra, Rishab
Narechania, Arpit
Endert, Alex
Stasko, John
Dou, Wenwen
author_facet Sah, Subham
Mitra, Rishab
Narechania, Arpit
Endert, Alex
Stasko, John
Dou, Wenwen
contents Recently, large language models (LLMs) have shown great promise in translating natural language (NL) queries into visualizations, but their "black-box" nature often limits explainability and debuggability. In response, we present a comprehensive text prompt that, given a tabular dataset and an NL query about the dataset, generates an analytic specification including (detected) data attributes, (inferred) analytic tasks, and (recommended) visualizations. This specification captures key aspects of the query translation process, affording both explainability and debuggability. For instance, it provides mappings from the detected entities to the corresponding phrases in the input query, as well as the specific visual design principles that determined the visualization recommendations. Moreover, unlike prior LLM-based approaches, our prompt supports conversational interaction and ambiguity detection capabilities. In this paper, we detail the iterative process of curating our prompt, present a preliminary performance evaluation using GPT-4, and discuss the strengths and limitations of LLMs at various stages of query translation. The prompt is open-source and integrated into NL4DV, a popular Python-based natural language toolkit for visualization, which can be accessed at https://nl4dv.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13391
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generating Analytic Specifications for Data Visualization from Natural Language Queries using Large Language Models
Sah, Subham
Mitra, Rishab
Narechania, Arpit
Endert, Alex
Stasko, John
Dou, Wenwen
Human-Computer Interaction
Recently, large language models (LLMs) have shown great promise in translating natural language (NL) queries into visualizations, but their "black-box" nature often limits explainability and debuggability. In response, we present a comprehensive text prompt that, given a tabular dataset and an NL query about the dataset, generates an analytic specification including (detected) data attributes, (inferred) analytic tasks, and (recommended) visualizations. This specification captures key aspects of the query translation process, affording both explainability and debuggability. For instance, it provides mappings from the detected entities to the corresponding phrases in the input query, as well as the specific visual design principles that determined the visualization recommendations. Moreover, unlike prior LLM-based approaches, our prompt supports conversational interaction and ambiguity detection capabilities. In this paper, we detail the iterative process of curating our prompt, present a preliminary performance evaluation using GPT-4, and discuss the strengths and limitations of LLMs at various stages of query translation. The prompt is open-source and integrated into NL4DV, a popular Python-based natural language toolkit for visualization, which can be accessed at https://nl4dv.github.io.
title Generating Analytic Specifications for Data Visualization from Natural Language Queries using Large Language Models
topic Human-Computer Interaction
url https://arxiv.org/abs/2408.13391