Insights into Natural Language Database Query Errors: From Attention Misalignment to User Handling Strategies

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ning, Zheng, Tian, Yuan, Zhang, Zheng, Zhang, Tianyi, Li, Toby
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911775200378880
author Ning, Zheng
Tian, Yuan
Zhang, Zheng
Zhang, Tianyi
Li, Toby
author_facet Ning, Zheng
Tian, Yuan
Zhang, Zheng
Zhang, Tianyi
Li, Toby
contents Querying structured databases with natural language (NL2SQL) has remained a difficult problem for years. Recently, the advancement of machine learning (ML), natural language processing (NLP), and large language models (LLM) have led to significant improvements in performance, with the best model achieving ~85% percent accuracy on the benchmark Spider dataset. However, there is a lack of a systematic understanding of the types, causes, and effectiveness of error-handling mechanisms of errors for erroneous queries nowadays. To bridge the gap, a taxonomy of errors made by four representative NL2SQL models was built in this work, along with an in-depth analysis of the errors. Second, the causes of model errors were explored by analyzing the model-human attention alignment to the natural language query. Last, a within-subjects user study with 26 participants was conducted to investigate the effectiveness of three interactive error-handling mechanisms in NL2SQL. Findings from this paper shed light on the design of model structure and error discovery and repair strategies for natural language data query interfaces in the future.
format Preprint
id arxiv_https___arxiv_org_abs_2402_07304
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Insights into Natural Language Database Query Errors: From Attention Misalignment to User Handling Strategies
Ning, Zheng
Tian, Yuan
Zhang, Zheng
Zhang, Tianyi
Li, Toby
Human-Computer Interaction
Querying structured databases with natural language (NL2SQL) has remained a difficult problem for years. Recently, the advancement of machine learning (ML), natural language processing (NLP), and large language models (LLM) have led to significant improvements in performance, with the best model achieving ~85% percent accuracy on the benchmark Spider dataset. However, there is a lack of a systematic understanding of the types, causes, and effectiveness of error-handling mechanisms of errors for erroneous queries nowadays. To bridge the gap, a taxonomy of errors made by four representative NL2SQL models was built in this work, along with an in-depth analysis of the errors. Second, the causes of model errors were explored by analyzing the model-human attention alignment to the natural language query. Last, a within-subjects user study with 26 participants was conducted to investigate the effectiveness of three interactive error-handling mechanisms in NL2SQL. Findings from this paper shed light on the design of model structure and error discovery and repair strategies for natural language data query interfaces in the future.
title Insights into Natural Language Database Query Errors: From Attention Misalignment to User Handling Strategies
topic Human-Computer Interaction
url https://arxiv.org/abs/2402.07304