SQL-Encoder: Improving NL2SQL In-Context Learning Through a Context-Aware Encoder

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pourreza, Mohammadreza, Rafiei, Davood, Feng, Yuxi, Li, Raymond, Fan, Zhenan, Zhang, Weiwei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914727435698176
author Pourreza, Mohammadreza
Rafiei, Davood
Feng, Yuxi
Li, Raymond
Fan, Zhenan
Zhang, Weiwei
author_facet Pourreza, Mohammadreza
Rafiei, Davood
Feng, Yuxi
Li, Raymond
Fan, Zhenan
Zhang, Weiwei
contents Detecting structural similarity between queries is essential for selecting examples in in-context learning models. However, assessing structural similarity based solely on the natural language expressions of queries, without considering SQL queries, presents a significant challenge. This paper explores the significance of this similarity metric and proposes a model for accurately estimating it. To achieve this, we leverage a dataset comprising 170k question pairs, meticulously curated to train a similarity prediction model. Our comprehensive evaluation demonstrates that the proposed model adeptly captures the structural similarity between questions, as evidenced by improvements in Kendall-Tau distance and precision@k metrics. Notably, our model outperforms strong competitive embedding models from OpenAI and Cohere. Furthermore, compared to these competitive models, our proposed encoder enhances the downstream performance of NL2SQL models in 1-shot in-context learning scenarios by 1-2\% for GPT-3.5-turbo, 4-8\% for CodeLlama-7B, and 2-3\% for CodeLlama-13B.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16204
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SQL-Encoder: Improving NL2SQL In-Context Learning Through a Context-Aware Encoder
Pourreza, Mohammadreza
Rafiei, Davood
Feng, Yuxi
Li, Raymond
Fan, Zhenan
Zhang, Weiwei
Computation and Language
Databases
Human-Computer Interaction
Detecting structural similarity between queries is essential for selecting examples in in-context learning models. However, assessing structural similarity based solely on the natural language expressions of queries, without considering SQL queries, presents a significant challenge. This paper explores the significance of this similarity metric and proposes a model for accurately estimating it. To achieve this, we leverage a dataset comprising 170k question pairs, meticulously curated to train a similarity prediction model. Our comprehensive evaluation demonstrates that the proposed model adeptly captures the structural similarity between questions, as evidenced by improvements in Kendall-Tau distance and precision@k metrics. Notably, our model outperforms strong competitive embedding models from OpenAI and Cohere. Furthermore, compared to these competitive models, our proposed encoder enhances the downstream performance of NL2SQL models in 1-shot in-context learning scenarios by 1-2\% for GPT-3.5-turbo, 4-8\% for CodeLlama-7B, and 2-3\% for CodeLlama-13B.
title SQL-Encoder: Improving NL2SQL In-Context Learning Through a Context-Aware Encoder
topic Computation and Language
Databases
Human-Computer Interaction
url https://arxiv.org/abs/2403.16204