Evolution of the "long tail" concept for scientific data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stahlman, Gretchen R., Kouper, Inna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910751245991936
author Stahlman, Gretchen R.
Kouper, Inna
author_facet Stahlman, Gretchen R.
Kouper, Inna
contents This review paper explores the evolution of discussions about "long-tail" scientific data in the scholarly literature. The "long-tail" concept, originally used to explain trends in digital consumer goods, was first applied to scientific data in 2007 to refer to a vast array of smaller, heterogeneous data collections that cumulatively represent a substantial portion of scientific knowledge. However, these datasets, often referred to as "long-tail data," are frequently mismanaged or overlooked due to inadequate data management practices and institutional support. This paper examines the changing landscape of discussions about long-tail data over time, situated within broader ecosystems of research data management and the natural interplay between "big" and "small" data. The review also bridges discussions on data curation in Library & Information Science (LIS) and domain-specific contexts, contributing to a more comprehensive understanding of the long-tail concept's utility for effective data management outcomes. The review aims to provide a more comprehensive understanding of this concept, its terminological diversity in the literature, and its utility for guiding data management, overall informing current and future information science research and practice.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13307
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evolution of the "long tail" concept for scientific data
Stahlman, Gretchen R.
Kouper, Inna
Digital Libraries
Social and Information Networks
This review paper explores the evolution of discussions about "long-tail" scientific data in the scholarly literature. The "long-tail" concept, originally used to explain trends in digital consumer goods, was first applied to scientific data in 2007 to refer to a vast array of smaller, heterogeneous data collections that cumulatively represent a substantial portion of scientific knowledge. However, these datasets, often referred to as "long-tail data," are frequently mismanaged or overlooked due to inadequate data management practices and institutional support. This paper examines the changing landscape of discussions about long-tail data over time, situated within broader ecosystems of research data management and the natural interplay between "big" and "small" data. The review also bridges discussions on data curation in Library & Information Science (LIS) and domain-specific contexts, contributing to a more comprehensive understanding of the long-tail concept's utility for effective data management outcomes. The review aims to provide a more comprehensive understanding of this concept, its terminological diversity in the literature, and its utility for guiding data management, overall informing current and future information science research and practice.
title Evolution of the "long tail" concept for scientific data
topic Digital Libraries
Social and Information Networks
url https://arxiv.org/abs/2412.13307