Semantically Aligned Question and Code Generation for Automated Insight Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Singha, Ananya, Chopra, Bhavya, Khatry, Anirudh, Gulwani, Sumit, Henley, Austin Z., Le, Vu, Parnin, Chris, Singh, Mukul, Verbruggen, Gust
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910432479936512
author Singha, Ananya
Chopra, Bhavya
Khatry, Anirudh
Gulwani, Sumit
Henley, Austin Z.
Le, Vu
Parnin, Chris
Singh, Mukul
Verbruggen, Gust
author_facet Singha, Ananya
Chopra, Bhavya
Khatry, Anirudh
Gulwani, Sumit
Henley, Austin Z.
Le, Vu
Parnin, Chris
Singh, Mukul
Verbruggen, Gust
contents Automated insight generation is a common tactic for helping knowledge workers, such as data scientists, to quickly understand the potential value of new and unfamiliar data. Unfortunately, automated insights produced by large-language models can generate code that does not correctly correspond (or align) to the insight. In this paper, we leverage the semantic knowledge of large language models to generate targeted and insightful questions about data and the corresponding code to answer those questions. Then through an empirical study on data from Open-WikiTable, we show that embeddings can be effectively used for filtering out semantically unaligned pairs of question and code. Additionally, we found that generating questions and code together yields more diverse questions.
format Preprint
id arxiv_https___arxiv_org_abs_2405_01556
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Semantically Aligned Question and Code Generation for Automated Insight Generation
Singha, Ananya
Chopra, Bhavya
Khatry, Anirudh
Gulwani, Sumit
Henley, Austin Z.
Le, Vu
Parnin, Chris
Singh, Mukul
Verbruggen, Gust
Software Engineering
Artificial Intelligence
Computation and Language
Automated insight generation is a common tactic for helping knowledge workers, such as data scientists, to quickly understand the potential value of new and unfamiliar data. Unfortunately, automated insights produced by large-language models can generate code that does not correctly correspond (or align) to the insight. In this paper, we leverage the semantic knowledge of large language models to generate targeted and insightful questions about data and the corresponding code to answer those questions. Then through an empirical study on data from Open-WikiTable, we show that embeddings can be effectively used for filtering out semantically unaligned pairs of question and code. Additionally, we found that generating questions and code together yields more diverse questions.
title Semantically Aligned Question and Code Generation for Automated Insight Generation
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.01556