A Datalake for Data-driven Social Science Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arya, Puneet, Sahasrabudhe, Ojas, Srivastav, Adwaiya, Das, Partha Pratim, Ramanath, Maya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912742791708672
author Arya, Puneet
Sahasrabudhe, Ojas
Srivastav, Adwaiya
Das, Partha Pratim
Ramanath, Maya
author_facet Arya, Puneet
Sahasrabudhe, Ojas
Srivastav, Adwaiya
Das, Partha Pratim
Ramanath, Maya
contents Social science research increasingly demands data-driven insights, yet researchers often face barriers such as lack of technical expertise, inconsistent data formats, and limited access to reliable datasets.Social science research increasingly demands data-driven insights, yet researchers often face barriers such as lack of technical expertise, inconsistent data formats, and limited access to reliable datasets. In this paper, we present a Datalake infrastructure tailored to the needs of interdisciplinary social science research. Our system supports ingestion and integration of diverse data types, automatic provenance and version tracking, role-based access control, and built-in tools for visualization and analysis. We demonstrate the utility of our Datalake using real-world use cases spanning governance, health, and education. A detailed walkthrough of one such use case -- analyzing the relationship between income, education, and infant mortality -- shows how our platform streamlines the research process while maintaining transparency and reproducibility. We argue that such infrastructure can democratize access to advanced data science practices, especially for NGOs, students, and grassroots organizations. The Datalake continues to evolve with plans to support ML pipelines, mobile access, and citizen data feedback mechanisms.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02463
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Datalake for Data-driven Social Science Research
Arya, Puneet
Sahasrabudhe, Ojas
Srivastav, Adwaiya
Das, Partha Pratim
Ramanath, Maya
Databases
Computers and Society
Social science research increasingly demands data-driven insights, yet researchers often face barriers such as lack of technical expertise, inconsistent data formats, and limited access to reliable datasets.Social science research increasingly demands data-driven insights, yet researchers often face barriers such as lack of technical expertise, inconsistent data formats, and limited access to reliable datasets. In this paper, we present a Datalake infrastructure tailored to the needs of interdisciplinary social science research. Our system supports ingestion and integration of diverse data types, automatic provenance and version tracking, role-based access control, and built-in tools for visualization and analysis. We demonstrate the utility of our Datalake using real-world use cases spanning governance, health, and education. A detailed walkthrough of one such use case -- analyzing the relationship between income, education, and infant mortality -- shows how our platform streamlines the research process while maintaining transparency and reproducibility. We argue that such infrastructure can democratize access to advanced data science practices, especially for NGOs, students, and grassroots organizations. The Datalake continues to evolve with plans to support ML pipelines, mobile access, and citizen data feedback mechanisms.
title A Datalake for Data-driven Social Science Research
topic Databases
Computers and Society
url https://arxiv.org/abs/2512.02463