Tursio Database Search: How far are we from ChatGPT?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jain, Sulbha, Tripathi, Shivani, Qiao, Shi, Jindal, Alekh
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912979911442432
author Jain, Sulbha
Tripathi, Shivani
Qiao, Shi
Jindal, Alekh
author_facet Jain, Sulbha
Tripathi, Shivani
Qiao, Shi
Jindal, Alekh
contents Business users need to search enterprise databases using natural language, just as they now search the web using ChatGPT or Perplexity. However, existing benchmarks -- designed for open-domain QA or text-to-SQL -- do not evaluate the end-to-end quality of such a search experience. We present an evaluation framework for structured database search that generates realistic banking queries across varying difficulty levels and assesses answer quality using relevance, safety, and conversational metrics via an LLM-as-judge approach. We apply this framework to compare Tursio, a database search platform, against ChatGPT and Perplexity on a credit union banking schema. Our results show that Tursio achieves answer relevancy statistically comparable to both baselines (97.8% vs. 98.1% on simple, 90.0% vs. 100.0% on medium, 89.5% vs. 100.0% on hard questions), even though Tursio answers from a structured database while the baselines generate responses from the open web. We analyze the failure modes, identify database completeness as the primary bottleneck, and outline directions for improving both the evaluation methodology and the systems under evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18835
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Tursio Database Search: How far are we from ChatGPT?
Jain, Sulbha
Tripathi, Shivani
Qiao, Shi
Jindal, Alekh
Databases
Business users need to search enterprise databases using natural language, just as they now search the web using ChatGPT or Perplexity. However, existing benchmarks -- designed for open-domain QA or text-to-SQL -- do not evaluate the end-to-end quality of such a search experience. We present an evaluation framework for structured database search that generates realistic banking queries across varying difficulty levels and assesses answer quality using relevance, safety, and conversational metrics via an LLM-as-judge approach. We apply this framework to compare Tursio, a database search platform, against ChatGPT and Perplexity on a credit union banking schema. Our results show that Tursio achieves answer relevancy statistically comparable to both baselines (97.8% vs. 98.1% on simple, 90.0% vs. 100.0% on medium, 89.5% vs. 100.0% on hard questions), even though Tursio answers from a structured database while the baselines generate responses from the open web. We analyze the failure modes, identify database completeness as the primary bottleneck, and outline directions for improving both the evaluation methodology and the systems under evaluation.
title Tursio Database Search: How far are we from ChatGPT?
topic Databases
url https://arxiv.org/abs/2603.18835