Natural Language Dataset Generation Framework for Visualizations Powered by Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ko, Hyung-Kwon, Jeon, Hyeon, Park, Gwanmo, Kim, Dae Hyun, Kim, Nam Wook, Kim, Juho, Seo, Jinwook
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910303984287744
author Ko, Hyung-Kwon
Jeon, Hyeon
Park, Gwanmo
Kim, Dae Hyun
Kim, Nam Wook
Kim, Juho
Seo, Jinwook
author_facet Ko, Hyung-Kwon
Jeon, Hyeon
Park, Gwanmo
Kim, Dae Hyun
Kim, Nam Wook
Kim, Juho
Seo, Jinwook
contents We introduce VL2NL, a Large Language Model (LLM) framework that generates rich and diverse NL datasets using only Vega-Lite specifications as input, thereby streamlining the development of Natural Language Interfaces (NLIs) for data visualization. To synthesize relevant chart semantics accurately and enhance syntactic diversity in each NL dataset, we leverage 1) a guided discovery incorporated into prompting so that LLMs can steer themselves to create faithful NL datasets in a self-directed manner; 2) a score-based paraphrasing to augment NL syntax along with four language axes. We also present a new collection of 1,981 real-world Vega-Lite specifications that have increased diversity and complexity than existing chart collections. When tested on our chart collection, VL2NL extracted chart semantics and generated L1/L2 captions with 89.4% and 76.0% accuracy, respectively. It also demonstrated generating and paraphrasing utterances and questions with greater diversity compared to the benchmarks. Last, we discuss how our NL datasets and framework can be utilized in real-world scenarios. The codes and chart collection are available at https://github.com/hyungkwonko/chart-llm.
format Preprint
id arxiv_https___arxiv_org_abs_2309_10245
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Natural Language Dataset Generation Framework for Visualizations Powered by Large Language Models
Ko, Hyung-Kwon
Jeon, Hyeon
Park, Gwanmo
Kim, Dae Hyun
Kim, Nam Wook
Kim, Juho
Seo, Jinwook
Human-Computer Interaction
We introduce VL2NL, a Large Language Model (LLM) framework that generates rich and diverse NL datasets using only Vega-Lite specifications as input, thereby streamlining the development of Natural Language Interfaces (NLIs) for data visualization. To synthesize relevant chart semantics accurately and enhance syntactic diversity in each NL dataset, we leverage 1) a guided discovery incorporated into prompting so that LLMs can steer themselves to create faithful NL datasets in a self-directed manner; 2) a score-based paraphrasing to augment NL syntax along with four language axes. We also present a new collection of 1,981 real-world Vega-Lite specifications that have increased diversity and complexity than existing chart collections. When tested on our chart collection, VL2NL extracted chart semantics and generated L1/L2 captions with 89.4% and 76.0% accuracy, respectively. It also demonstrated generating and paraphrasing utterances and questions with greater diversity compared to the benchmarks. Last, we discuss how our NL datasets and framework can be utilized in real-world scenarios. The codes and chart collection are available at https://github.com/hyungkwonko/chart-llm.
title Natural Language Dataset Generation Framework for Visualizations Powered by Large Language Models
topic Human-Computer Interaction
url https://arxiv.org/abs/2309.10245