Text-to-SQL Oriented to the Process Mining Domain: A PT-EN Dataset for Query Translation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911151093186560 |
|---|---|
| author | Yamate, Bruno Yui Neubauer, Thais Rodrigues Fantinato, Marcelo Peres, Sarajane Marques |
| author_facet | Yamate, Bruno Yui Neubauer, Thais Rodrigues Fantinato, Marcelo Peres, Sarajane Marques |
| contents | This paper introduces text-2-SQL-4-PM, a bilingual (Portuguese-English) benchmark dataset designed for the text-to-SQL task in the process mining domain. Text-to-SQL conversion facilitates natural language querying of databases, increasing accessibility for users without SQL expertise and productivity for those that are experts. The text-2-SQL-4-PM dataset is customized to address the unique challenges of process mining, including specialized vocabularies and single-table relational structures derived from event logs. The dataset comprises 1,655 natural language utterances, including human-generated paraphrases, 205 SQL statements, and ten qualifiers. Methods include manual curation by experts, professional translations, and a detailed annotation process to enable nuanced analyses of task complexity. Additionally, a baseline study using GPT-3.5 Turbo demonstrates the feasibility and utility of the dataset for text-to-SQL applications. The results show that text-2-SQL-4-PM supports evaluation of text-to-SQL implementations, offering broader applicability for semantic parsing and other natural language processing tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_09684 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Text-to-SQL Oriented to the Process Mining Domain: A PT-EN Dataset for Query Translation Yamate, Bruno Yui Neubauer, Thais Rodrigues Fantinato, Marcelo Peres, Sarajane Marques Information Retrieval Artificial Intelligence Computation and Language Databases This paper introduces text-2-SQL-4-PM, a bilingual (Portuguese-English) benchmark dataset designed for the text-to-SQL task in the process mining domain. Text-to-SQL conversion facilitates natural language querying of databases, increasing accessibility for users without SQL expertise and productivity for those that are experts. The text-2-SQL-4-PM dataset is customized to address the unique challenges of process mining, including specialized vocabularies and single-table relational structures derived from event logs. The dataset comprises 1,655 natural language utterances, including human-generated paraphrases, 205 SQL statements, and ten qualifiers. Methods include manual curation by experts, professional translations, and a detailed annotation process to enable nuanced analyses of task complexity. Additionally, a baseline study using GPT-3.5 Turbo demonstrates the feasibility and utility of the dataset for text-to-SQL applications. The results show that text-2-SQL-4-PM supports evaluation of text-to-SQL implementations, offering broader applicability for semantic parsing and other natural language processing tasks. |
| title | Text-to-SQL Oriented to the Process Mining Domain: A PT-EN Dataset for Query Translation |
| topic | Information Retrieval Artificial Intelligence Computation and Language Databases |
| url | https://arxiv.org/abs/2509.09684 |