Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Singh, Jyotika, Sun, Weiyi, Agarwal, Amit, Krishnamurthy, Viji, Benajiba, Yassine, Ravi, Sujith, Roth, Dan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909873216684032
author Singh, Jyotika
Sun, Weiyi
Agarwal, Amit
Krishnamurthy, Viji
Benajiba, Yassine
Ravi, Sujith
Roth, Dan
author_facet Singh, Jyotika
Sun, Weiyi
Agarwal, Amit
Krishnamurthy, Viji
Benajiba, Yassine
Ravi, Sujith
Roth, Dan
contents In modern industry systems like multi-turn chat agents, Text-to-SQL technology bridges natural language (NL) questions and database (DB) querying. The conversion of tabular DB results into NL representations (NLRs) enables the chat-based interaction. Currently, NLR generation is typically handled by large language models (LLMs), but information loss or errors in presenting tabular results in NL remains largely unexplored. This paper introduces a novel evaluation method - Combo-Eval - for judgment of LLM-generated NLRs that combines the benefits of multiple existing methods, optimizing evaluation fidelity and achieving a significant reduction in LLM calls by 25-61%. Accompanying our method is NLR-BIRD, the first dedicated dataset for NLR benchmarking. Through human evaluations, we demonstrate the superior alignment of Combo-Eval with human judgments, applicable across scenarios with and without ground truth references.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23854
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs
Singh, Jyotika
Sun, Weiyi
Agarwal, Amit
Krishnamurthy, Viji
Benajiba, Yassine
Ravi, Sujith
Roth, Dan
Computation and Language
Artificial Intelligence
In modern industry systems like multi-turn chat agents, Text-to-SQL technology bridges natural language (NL) questions and database (DB) querying. The conversion of tabular DB results into NL representations (NLRs) enables the chat-based interaction. Currently, NLR generation is typically handled by large language models (LLMs), but information loss or errors in presenting tabular results in NL remains largely unexplored. This paper introduces a novel evaluation method - Combo-Eval - for judgment of LLM-generated NLRs that combines the benefits of multiple existing methods, optimizing evaluation fidelity and achieving a significant reduction in LLM calls by 25-61%. Accompanying our method is NLR-BIRD, the first dedicated dataset for NLR benchmarking. Through human evaluations, we demonstrate the superior alignment of Combo-Eval with human judgments, applicable across scenarios with and without ground truth references.
title Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.23854