Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Terrance, Wang, Shuyi, Preotiuc-Pietro, Daniel, Chandarana, Yash, Gupta, Chirag
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2505.23804
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911158756179968
author Liu, Terrance
Wang, Shuyi
Preotiuc-Pietro, Daniel
Chandarana, Yash
Gupta, Chirag
author_facet Liu, Terrance
Wang, Shuyi
Preotiuc-Pietro, Daniel
Chandarana, Yash
Gupta, Chirag
contents While large language models (LLMs) achieve strong performance on text-to-SQL parsing, they sometimes exhibit unexpected failures in which they are confidently incorrect. Building trustworthy text-to-SQL systems thus requires eliciting reliable uncertainty measures from the LLM. In this paper, we study the problem of providing a calibrated confidence score that conveys the likelihood of an output query being correct. Our work is the first to establish a benchmark for post-hoc calibration of LLM-based text-to-SQL parsing. In particular, we show that Platt scaling, a canonical method for calibration, provides substantial improvements over directly using raw model output probabilities as confidence scores. Furthermore, we propose a method for text-to-SQL calibration that leverages the structured nature of SQL queries to provide more granular signals of correctness, named "sub-clause frequency" (SCF) scores. Using multivariate Platt scaling (MPS), our extension of the canonical Platt scaling technique, we combine individual SCF scores into an overall accurate and calibrated score. Empirical evaluation on two popular text-to-SQL datasets shows that our approach of combining MPS and SCF yields further improvements in calibration and the related task of error detection over traditional Platt scaling.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23804
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequencies
Liu, Terrance
Wang, Shuyi
Preotiuc-Pietro, Daniel
Chandarana, Yash
Gupta, Chirag
Computation and Language
Artificial Intelligence
Machine Learning
While large language models (LLMs) achieve strong performance on text-to-SQL parsing, they sometimes exhibit unexpected failures in which they are confidently incorrect. Building trustworthy text-to-SQL systems thus requires eliciting reliable uncertainty measures from the LLM. In this paper, we study the problem of providing a calibrated confidence score that conveys the likelihood of an output query being correct. Our work is the first to establish a benchmark for post-hoc calibration of LLM-based text-to-SQL parsing. In particular, we show that Platt scaling, a canonical method for calibration, provides substantial improvements over directly using raw model output probabilities as confidence scores. Furthermore, we propose a method for text-to-SQL calibration that leverages the structured nature of SQL queries to provide more granular signals of correctness, named "sub-clause frequency" (SCF) scores. Using multivariate Platt scaling (MPS), our extension of the canonical Platt scaling technique, we combine individual SCF scores into an overall accurate and calibrated score. Empirical evaluation on two popular text-to-SQL datasets shows that our approach of combining MPS and SCF yields further improvements in calibration and the related task of error detection over traditional Platt scaling.
title Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequencies
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.23804