Database Entity Recognition with Data Augmentation and Deep Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Fu, Zikun, Yang, Chen, Davoudi, Kourosh, Pu, Ken Q.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916920472633344
author Fu, Zikun
Yang, Chen
Davoudi, Kourosh
Pu, Ken Q.
author_facet Fu, Zikun
Yang, Chen
Davoudi, Kourosh
Pu, Ken Q.
contents This paper addresses the challenge of Database Entity Recognition (DB-ER) in Natural Language Queries (NLQ). We present several key contributions to advance this field: (1) a human-annotated benchmark for DB-ER task, derived from popular text-to-sql benchmarks, (2) a novel data augmentation procedure that leverages automatic annotation of NLQs based on the corresponding SQL queries which are available in popular text-to-SQL benchmarks, (3) a specialized language model based entity recognition model using T5 as a backbone and two down-stream DB-ER tasks: sequence tagging and token classification for fine-tuning of backend and performing DB-ER respectively. We compared our DB-ER tagger with two state-of-the-art NER taggers, and observed better performance in both precision and recall for our model. The ablation evaluation shows that data augmentation boosts precision and recall by over 10%, while fine-tuning of the T5 backbone boosts these metrics by 5-10%.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19372
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Database Entity Recognition with Data Augmentation and Deep Learning
Fu, Zikun
Yang, Chen
Davoudi, Kourosh
Pu, Ken Q.
Computation and Language
Artificial Intelligence
Databases
Machine Learning
This paper addresses the challenge of Database Entity Recognition (DB-ER) in Natural Language Queries (NLQ). We present several key contributions to advance this field: (1) a human-annotated benchmark for DB-ER task, derived from popular text-to-sql benchmarks, (2) a novel data augmentation procedure that leverages automatic annotation of NLQs based on the corresponding SQL queries which are available in popular text-to-SQL benchmarks, (3) a specialized language model based entity recognition model using T5 as a backbone and two down-stream DB-ER tasks: sequence tagging and token classification for fine-tuning of backend and performing DB-ER respectively. We compared our DB-ER tagger with two state-of-the-art NER taggers, and observed better performance in both precision and recall for our model. The ablation evaluation shows that data augmentation boosts precision and recall by over 10%, while fine-tuning of the T5 backbone boosts these metrics by 5-10%.
title Database Entity Recognition with Data Augmentation and Deep Learning
topic Computation and Language
Artificial Intelligence
Databases
Machine Learning
url https://arxiv.org/abs/2508.19372