Saved in:
Bibliographic Details
Main Authors: Lang, Logan, Hernandez, Eduardo, Choudhary, Kamal, Romero, Aldo H.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.05311
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908331005706240
author Lang, Logan
Hernandez, Eduardo
Choudhary, Kamal
Romero, Aldo H.
author_facet Lang, Logan
Hernandez, Eduardo
Choudhary, Kamal
Romero, Aldo H.
contents Traditional data storage formats and databases often introduce complexities and inefficiencies that hinder rapid iteration and adaptability. To address these challenges, we introduce ParquetDB, a Python-based database framework that leverages the Parquet file format's optimized columnar storage. ParquetDB offers efficient serialization and deserialization, native support for complex and nested data types, reduced dependency on indexing through predicate pushdown filtering, and enhanced portability due to its file-based storage system. Benchmarks show that ParquetDB outperforms traditional databases like SQLite and MongoDB in managing large volumes of data, especially when using data formats compatible with PyArrow. We validate ParquetDB's practical utility by applying it to the Alexandria 3D Materials Database, efficiently handling approximately 4.8 million complex and nested records. By addressing the inherent limitations of existing data storage systems and continuously evolving to meet future demands, ParquetDB has the potential to significantly streamline data management processes and accelerate research development in data-driven fields.
format Preprint
id arxiv_https___arxiv_org_abs_2502_05311
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ParquetDB: A Lightweight Python Parquet-Based Database
Lang, Logan
Hernandez, Eduardo
Choudhary, Kamal
Romero, Aldo H.
Databases
Data Analysis, Statistics and Probability
Traditional data storage formats and databases often introduce complexities and inefficiencies that hinder rapid iteration and adaptability. To address these challenges, we introduce ParquetDB, a Python-based database framework that leverages the Parquet file format's optimized columnar storage. ParquetDB offers efficient serialization and deserialization, native support for complex and nested data types, reduced dependency on indexing through predicate pushdown filtering, and enhanced portability due to its file-based storage system. Benchmarks show that ParquetDB outperforms traditional databases like SQLite and MongoDB in managing large volumes of data, especially when using data formats compatible with PyArrow. We validate ParquetDB's practical utility by applying it to the Alexandria 3D Materials Database, efficiently handling approximately 4.8 million complex and nested records. By addressing the inherent limitations of existing data storage systems and continuously evolving to meet future demands, ParquetDB has the potential to significantly streamline data management processes and accelerate research development in data-driven fields.
title ParquetDB: A Lightweight Python Parquet-Based Database
topic Databases
Data Analysis, Statistics and Probability
url https://arxiv.org/abs/2502.05311