Saved in:
Bibliographic Details
Main Authors: Papicchio, Simone, Rossi, Simone, Cagliero, Luca, Papotti, Paolo
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.15077
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917455273656320
author Papicchio, Simone
Rossi, Simone
Cagliero, Luca
Papotti, Paolo
author_facet Papicchio, Simone
Rossi, Simone
Cagliero, Luca
Papotti, Paolo
contents Large Language Models (LLMs) can translate natural language into SQL, but small models struggle with multi-table and complex queries in Zero-Shot Learning (ZSL) settings. While Supervised Fine-Tuning (SFT) helps, it falls short for harder cases. To address this, we study how different reasoning strategies (general-purpose reasoning in ZSL, reasoning traces in SFT, and Reinforcement Learning with Verifiable Reward (RLVR) with novel reward functions) affect Text2SQL performance across four benchmarks. We show that partial scoring rewards, computed via SQL execution, are crucial for guiding models even when outputs are not fully correct. These fine-grained signals lead to consistently better Text2SQL outcomes. Small LLMs benefit most from reasoning-aware SFT and RL, with the 14B Qwen-Coder-2.5 surpassing 400B+ models on challenging datasets like BIRD.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15077
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL
Papicchio, Simone
Rossi, Simone
Cagliero, Luca
Papotti, Paolo
Machine Learning
Databases
Large Language Models (LLMs) can translate natural language into SQL, but small models struggle with multi-table and complex queries in Zero-Shot Learning (ZSL) settings. While Supervised Fine-Tuning (SFT) helps, it falls short for harder cases. To address this, we study how different reasoning strategies (general-purpose reasoning in ZSL, reasoning traces in SFT, and Reinforcement Learning with Verifiable Reward (RLVR) with novel reward functions) affect Text2SQL performance across four benchmarks. We show that partial scoring rewards, computed via SQL execution, are crucial for guiding models even when outputs are not fully correct. These fine-grained signals lead to consistently better Text2SQL outcomes. Small LLMs benefit most from reasoning-aware SFT and RL, with the 14B Qwen-Coder-2.5 surpassing 400B+ models on challenging datasets like BIRD.
title Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL
topic Machine Learning
Databases
url https://arxiv.org/abs/2504.15077