SchemaDB: Structures in Relational Datasets

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Christopher, Cody James, Moore, Kristen, Liebowitz, David
Formato: Preprint
Publicado: 2021
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913864487010304
author Christopher, Cody James
Moore, Kristen
Liebowitz, David
author_facet Christopher, Cody James
Moore, Kristen
Liebowitz, David
contents In this paper we introduce the SchemaDB data-set; a collection of relational database schemata in both sql and graph formats. Databases are not commonly shared publicly for reasons of privacy and security, so schemata are not available for study. Consequently, an understanding of database structures in the wild is lacking, and most examples found publicly belong to common development frameworks or are derived from textbooks or engine benchmark designs. SchemaDB contains 2,500 samples of relational schemata found in public repositories which we have standardised to MySQL syntax. We provide our gathering and transformation methodology, summary statistics, and structural analysis, and discuss potential downstream research tasks in several domains.
format Preprint
id arxiv_https___arxiv_org_abs_2111_12835
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle SchemaDB: Structures in Relational Datasets
Christopher, Cody James
Moore, Kristen
Liebowitz, David
Databases
Machine Learning
In this paper we introduce the SchemaDB data-set; a collection of relational database schemata in both sql and graph formats. Databases are not commonly shared publicly for reasons of privacy and security, so schemata are not available for study. Consequently, an understanding of database structures in the wild is lacking, and most examples found publicly belong to common development frameworks or are derived from textbooks or engine benchmark designs. SchemaDB contains 2,500 samples of relational schemata found in public repositories which we have standardised to MySQL syntax. We provide our gathering and transformation methodology, summary statistics, and structural analysis, and discuss potential downstream research tasks in several domains.
title SchemaDB: Structures in Relational Datasets
topic Databases
Machine Learning
url https://arxiv.org/abs/2111.12835