Declarative Techniques for NL Queries over Heterogeneous Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khabiri, Elham, Kephart, Jeffrey O., Heath III, Fenno F., Jayaraman, Srideepika, Tipu, Fateh A., Li, Yingjie, Shah, Dhruv, Fokoue, Achille, Bhamidipaty, Anu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909856072466432
author Khabiri, Elham
Kephart, Jeffrey O.
Heath III, Fenno F.
Jayaraman, Srideepika
Tipu, Fateh A.
Li, Yingjie
Shah, Dhruv
Fokoue, Achille
Bhamidipaty, Anu
author_facet Khabiri, Elham
Kephart, Jeffrey O.
Heath III, Fenno F.
Jayaraman, Srideepika
Tipu, Fateh A.
Li, Yingjie
Shah, Dhruv
Fokoue, Achille
Bhamidipaty, Anu
contents In many industrial settings, users wish to ask questions in natural language, the answers to which require assembling information from diverse structured data sources. With the advent of Large Language Models (LLMs), applications can now translate natural language questions into a set of API calls or database calls, execute them, and combine the results into an appropriate natural language response. However, these applications remain impractical in realistic industrial settings because they do not cope with the data source heterogeneity that typifies such environments. In this work, we simulate the heterogeneity of real industry settings by introducing two extensions of the popular Spider benchmark dataset that require a combination of database and API calls. Then, we introduce a declarative approach to handling such data heterogeneity and demonstrate that it copes with data source heterogeneity significantly better than state-of-the-art LLM-based agentic or imperative code generation systems. Our augmented benchmarks are available to the research community.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16470
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Declarative Techniques for NL Queries over Heterogeneous Data
Khabiri, Elham
Kephart, Jeffrey O.
Heath III, Fenno F.
Jayaraman, Srideepika
Tipu, Fateh A.
Li, Yingjie
Shah, Dhruv
Fokoue, Achille
Bhamidipaty, Anu
Databases
Artificial Intelligence
Software Engineering
In many industrial settings, users wish to ask questions in natural language, the answers to which require assembling information from diverse structured data sources. With the advent of Large Language Models (LLMs), applications can now translate natural language questions into a set of API calls or database calls, execute them, and combine the results into an appropriate natural language response. However, these applications remain impractical in realistic industrial settings because they do not cope with the data source heterogeneity that typifies such environments. In this work, we simulate the heterogeneity of real industry settings by introducing two extensions of the popular Spider benchmark dataset that require a combination of database and API calls. Then, we introduce a declarative approach to handling such data heterogeneity and demonstrate that it copes with data source heterogeneity significantly better than state-of-the-art LLM-based agentic or imperative code generation systems. Our augmented benchmarks are available to the research community.
title Declarative Techniques for NL Queries over Heterogeneous Data
topic Databases
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2510.16470