Confidence Estimation for Error Detection in Text-to-SQL Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Somov, Oleg, Tutubalina, Elena |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
by: Alekseev, Artem, et al.
Published: (2025)
by: Alekseev, Artem, et al.
Published: (2025)
Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards
by: Zhang, Yuxin, et al.
Published: (2025)
by: Zhang, Yuxin, et al.
Published: (2025)
Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation
by: Ranaldi, Federico, et al.
Published: (2024)
by: Ranaldi, Federico, et al.
Published: (2024)
Know When You're Wrong: Aligning Confidence with Correctness for LLM Error Detection
by: Xiaohu, Xie, et al.
Published: (2026)
by: Xiaohu, Xie, et al.
Published: (2026)
RingSQL: Generating Synthetic Data with Schema-Independent Templates for Text-to-SQL Reasoning Models
by: Sterbentz, Marko, et al.
Published: (2026)
by: Sterbentz, Marko, et al.
Published: (2026)
BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases
by: Koretsky, Mathew J., et al.
Published: (2025)
by: Koretsky, Mathew J., et al.
Published: (2025)
BookSQL: A Large Scale Text-to-SQL Dataset for Accounting Domain
by: Kumar, Rahul, et al.
Published: (2024)
by: Kumar, Rahul, et al.
Published: (2024)
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
by: Seleznyov, Mikhail, et al.
Published: (2025)
by: Seleznyov, Mikhail, et al.
Published: (2025)
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
by: Seleznyov, Mikhail, et al.
Published: (2026)
by: Seleznyov, Mikhail, et al.
Published: (2026)
Table-to-Text Generation with Pretrained Diffusion Models
by: Krylov, Aleksei S., et al.
Published: (2024)
by: Krylov, Aleksei S., et al.
Published: (2024)
Confidence Preservation Property in Knowledge Distillation Abstractions
by: Vengertsev, Dmitry, et al.
Published: (2024)
by: Vengertsev, Dmitry, et al.
Published: (2024)
SQLformer: Deep Auto-Regressive Query Graph Generation for Text-to-SQL Translation
by: Bazaga, Adrián, et al.
Published: (2023)
by: Bazaga, Adrián, et al.
Published: (2023)
nach0-pc: Multi-task Language Model with Molecular Point Cloud Encoder
by: Kuznetsov, Maksim, et al.
Published: (2024)
by: Kuznetsov, Maksim, et al.
Published: (2024)
Knowledge Base Construction for Knowledge-Augmented Text-to-SQL
by: Baek, Jinheon, et al.
Published: (2025)
by: Baek, Jinheon, et al.
Published: (2025)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
by: Mahaut, Matéo, et al.
Published: (2024)
by: Mahaut, Matéo, et al.
Published: (2024)
CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQL
by: Pourreza, Mohammadreza, et al.
Published: (2024)
by: Pourreza, Mohammadreza, et al.
Published: (2024)
Cheaper, Better, Faster, Stronger: Robust Text-to-SQL without Chain-of-Thought or Fine-Tuning
by: Dönder, Yusuf Denizay, et al.
Published: (2025)
by: Dönder, Yusuf Denizay, et al.
Published: (2025)
SQL-GEN: Bridging the Dialect Gap for Text-to-SQL Via Synthetic Data And Model Merging
by: Pourreza, Mohammadreza, et al.
Published: (2024)
by: Pourreza, Mohammadreza, et al.
Published: (2024)
Confidence Estimation for Text-to-SQL in Large Language Models
by: Maleki, Sepideh Entezari, et al.
Published: (2025)
by: Maleki, Sepideh Entezari, et al.
Published: (2025)
Semantic Captioning: Benchmark Dataset and Graph-Aware Few-Shot In-Context Learning for SQL2Text
by: Al-Lawati, Ali, et al.
Published: (2025)
by: Al-Lawati, Ali, et al.
Published: (2025)
Using Source-Side Confidence Estimation for Reliable Translation into Unfamiliar Languages
by: Sible, Kenneth J., et al.
Published: (2025)
by: Sible, Kenneth J., et al.
Published: (2025)
Self-Evaluating LLMs for Multi-Step Tasks: Stepwise Confidence Estimation for Failure Detection
by: Mavi, Vaibhav, et al.
Published: (2025)
by: Mavi, Vaibhav, et al.
Published: (2025)
Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequencies
by: Liu, Terrance, et al.
Published: (2025)
by: Liu, Terrance, et al.
Published: (2025)
RASL: Retrieval Augmented Schema Linking for Massive Database Text-to-SQL
by: Eben, Jeffrey, et al.
Published: (2025)
by: Eben, Jeffrey, et al.
Published: (2025)
On the Security Vulnerabilities of Text-to-SQL Models
by: Peng, Xutan, et al.
Published: (2022)
by: Peng, Xutan, et al.
Published: (2022)
SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas
by: Wolff, Cornelius, et al.
Published: (2025)
by: Wolff, Cornelius, et al.
Published: (2025)
DISCERN: Decoding Systematic Errors in Natural Language for Text Classifiers
by: Menon, Rakesh R., et al.
Published: (2024)
by: Menon, Rakesh R., et al.
Published: (2024)
A Context-Aware Dual-Metric Framework for Confidence Estimation in Large Language Models
by: Yuan, Mingruo, et al.
Published: (2025)
by: Yuan, Mingruo, et al.
Published: (2025)
Trust in One Round: Confidence Estimation for Large Language Models via Structural Signals
by: Yang, Pengyue, et al.
Published: (2026)
by: Yang, Pengyue, et al.
Published: (2026)
Show Your Work with Confidence: Confidence Bands for Tuning Curves
by: Lourie, Nicholas, et al.
Published: (2023)
by: Lourie, Nicholas, et al.
Published: (2023)
Confidence Regularized Masked Language Modeling using Text Length
by: Ji, Seunghyun, et al.
Published: (2025)
by: Ji, Seunghyun, et al.
Published: (2025)
FLEX: Expert-level False-Less EXecution Metric for Reliable Text-to-SQL Benchmark
by: Kim, Heegyu, et al.
Published: (2024)
by: Kim, Heegyu, et al.
Published: (2024)
EvoSchema: Towards Text-to-SQL Robustness Against Schema Evolution
by: Zhang, Tianshu, et al.
Published: (2026)
by: Zhang, Tianshu, et al.
Published: (2026)
End-to-end Text-to-SQL Generation within an Analytics Insight Engine
by: Maamari, Karime, et al.
Published: (2024)
by: Maamari, Karime, et al.
Published: (2024)
Pareto Optimal Learning for Estimating Large Language Model Errors
by: Zhao, Theodore, et al.
Published: (2023)
by: Zhao, Theodore, et al.
Published: (2023)
Fact-Consistency Evaluation of Text-to-SQL Generation for Business Intelligence Using Exaone 3.5
by: Choi, Jeho
Published: (2025)
by: Choi, Jeho
Published: (2025)
H-STAR: LLM-driven Hybrid SQL-Text Adaptive Reasoning on Tables
by: Abhyankar, Nikhil, et al.
Published: (2024)
by: Abhyankar, Nikhil, et al.
Published: (2024)
Human Texts Are Outliers: Detecting LLM-generated Texts via Out-of-distribution Detection
by: Zeng, Cong, et al.
Published: (2025)
by: Zeng, Cong, et al.
Published: (2025)
Confidence over Time: Confidence Calibration with Temporal Logic for Large Language Model Reasoning
by: Mao, Zhenjiang, et al.
Published: (2026)
by: Mao, Zhenjiang, et al.
Published: (2026)
Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
by: Elgabry, Menna, et al.
Published: (2025)
by: Elgabry, Menna, et al.
Published: (2025)
Similar Items
-
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
by: Alekseev, Artem, et al.
Published: (2025) -
Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards
by: Zhang, Yuxin, et al.
Published: (2025) -
Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation
by: Ranaldi, Federico, et al.
Published: (2024) -
Know When You're Wrong: Aligning Confidence with Correctness for LLM Error Detection
by: Xiaohu, Xie, et al.
Published: (2026) -
RingSQL: Generating Synthetic Data with Schema-Independent Templates for Text-to-SQL Reasoning Models
by: Sterbentz, Marko, et al.
Published: (2026)