SQL-Encoder: Improving NL2SQL In-Context Learning Through a Context-Aware Encoder
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914727435698176 |
|---|---|
| author | Pourreza, Mohammadreza Rafiei, Davood Feng, Yuxi Li, Raymond Fan, Zhenan Zhang, Weiwei |
| author_facet | Pourreza, Mohammadreza Rafiei, Davood Feng, Yuxi Li, Raymond Fan, Zhenan Zhang, Weiwei |
| contents | Detecting structural similarity between queries is essential for selecting examples in in-context learning models. However, assessing structural similarity based solely on the natural language expressions of queries, without considering SQL queries, presents a significant challenge. This paper explores the significance of this similarity metric and proposes a model for accurately estimating it. To achieve this, we leverage a dataset comprising 170k question pairs, meticulously curated to train a similarity prediction model. Our comprehensive evaluation demonstrates that the proposed model adeptly captures the structural similarity between questions, as evidenced by improvements in Kendall-Tau distance and precision@k metrics. Notably, our model outperforms strong competitive embedding models from OpenAI and Cohere. Furthermore, compared to these competitive models, our proposed encoder enhances the downstream performance of NL2SQL models in 1-shot in-context learning scenarios by 1-2\% for GPT-3.5-turbo, 4-8\% for CodeLlama-7B, and 2-3\% for CodeLlama-13B. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_16204 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SQL-Encoder: Improving NL2SQL In-Context Learning Through a Context-Aware Encoder Pourreza, Mohammadreza Rafiei, Davood Feng, Yuxi Li, Raymond Fan, Zhenan Zhang, Weiwei Computation and Language Databases Human-Computer Interaction Detecting structural similarity between queries is essential for selecting examples in in-context learning models. However, assessing structural similarity based solely on the natural language expressions of queries, without considering SQL queries, presents a significant challenge. This paper explores the significance of this similarity metric and proposes a model for accurately estimating it. To achieve this, we leverage a dataset comprising 170k question pairs, meticulously curated to train a similarity prediction model. Our comprehensive evaluation demonstrates that the proposed model adeptly captures the structural similarity between questions, as evidenced by improvements in Kendall-Tau distance and precision@k metrics. Notably, our model outperforms strong competitive embedding models from OpenAI and Cohere. Furthermore, compared to these competitive models, our proposed encoder enhances the downstream performance of NL2SQL models in 1-shot in-context learning scenarios by 1-2\% for GPT-3.5-turbo, 4-8\% for CodeLlama-7B, and 2-3\% for CodeLlama-13B. |
| title | SQL-Encoder: Improving NL2SQL In-Context Learning Through a Context-Aware Encoder |
| topic | Computation and Language Databases Human-Computer Interaction |
| url | https://arxiv.org/abs/2403.16204 |