Text2Schema: Filling the Gap in Designing Database Table Structures based on Natural Language

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Qin, Li, Youhuan, Feng, Yansong, Chen, Si, Li, Ziming, Zhang, Pan, Si, Zihui, Chen, Yixuan, Shi, Zhichao, Huang, Zebin, Chen, Guo, Jin, Wenqiang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908598257319936
author Wang, Qin
Li, Youhuan
Feng, Yansong
Chen, Si
Li, Ziming
Zhang, Pan
Si, Zihui
Chen, Yixuan
Shi, Zhichao
Huang, Zebin
Chen, Guo
Jin, Wenqiang
author_facet Wang, Qin
Li, Youhuan
Feng, Yansong
Chen, Si
Li, Ziming
Zhang, Pan
Si, Zihui
Chen, Yixuan
Shi, Zhichao
Huang, Zebin
Chen, Guo
Jin, Wenqiang
contents People without a database background usually rely on file systems or tools such as Excel for data management, which often lead to redundancy and data inconsistency. Relational databases possess strong data management capabilities, but require a high level of professional expertise from users. Although there are already many works on Text2SQL to automate the translation of natural language into SQL queries for data manipulation, all of them presuppose that the database schema is pre-designed. In practice, schema design itself demands domain expertise, and research on directly generating schemas from textual requirements remains unexplored. In this paper, we systematically define a new problem, called Text2Schema, to convert a natural language text requirement into a relational database schema. With an effective Text2Schema technique, users can effortlessly create database table structures using natural language, and subsequently leverage existing Text2SQL techniques to perform data manipulations, which significantly narrows the gap between non-technical personnel and highly efficient, versatile relational database systems. We propose SchemaAgent, an LLM-based multi-agent framework for Text2Schema. We emulate the workflow of manual schema design by assigning specialized roles to agents and enabling effective collaboration to refine their respective subtasks. We also incorporate dedicated roles for reflection and inspection, along with an innovative error detection and correction mechanism to identify and rectify issues across various phases. Moreover, we build and open source a benchmark containing 381 pairs of requirement description and schema. Experimental results demonstrate the superiority of our approach over comparative work.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23886
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text2Schema: Filling the Gap in Designing Database Table Structures based on Natural Language
Wang, Qin
Li, Youhuan
Feng, Yansong
Chen, Si
Li, Ziming
Zhang, Pan
Si, Zihui
Chen, Yixuan
Shi, Zhichao
Huang, Zebin
Chen, Guo
Jin, Wenqiang
Databases
Artificial Intelligence
People without a database background usually rely on file systems or tools such as Excel for data management, which often lead to redundancy and data inconsistency. Relational databases possess strong data management capabilities, but require a high level of professional expertise from users. Although there are already many works on Text2SQL to automate the translation of natural language into SQL queries for data manipulation, all of them presuppose that the database schema is pre-designed. In practice, schema design itself demands domain expertise, and research on directly generating schemas from textual requirements remains unexplored. In this paper, we systematically define a new problem, called Text2Schema, to convert a natural language text requirement into a relational database schema. With an effective Text2Schema technique, users can effortlessly create database table structures using natural language, and subsequently leverage existing Text2SQL techniques to perform data manipulations, which significantly narrows the gap between non-technical personnel and highly efficient, versatile relational database systems. We propose SchemaAgent, an LLM-based multi-agent framework for Text2Schema. We emulate the workflow of manual schema design by assigning specialized roles to agents and enabling effective collaboration to refine their respective subtasks. We also incorporate dedicated roles for reflection and inspection, along with an innovative error detection and correction mechanism to identify and rectify issues across various phases. Moreover, we build and open source a benchmark containing 381 pairs of requirement description and schema. Experimental results demonstrate the superiority of our approach over comparative work.
title Text2Schema: Filling the Gap in Designing Database Table Structures based on Natural Language
topic Databases
Artificial Intelligence
url https://arxiv.org/abs/2503.23886