GRIN Transfer: A production-ready tool for libraries to retrieve digital copies from Google Books

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Daly, Liza, Cargnelutti, Matteo, Brobston, Catherine, Hess, John, Leppert, Greg, Watson, Amanda, Zittrain, Jonathan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915622166724608
author Daly, Liza
Cargnelutti, Matteo
Brobston, Catherine
Hess, John
Leppert, Greg
Watson, Amanda
Zittrain, Jonathan
author_facet Daly, Liza
Cargnelutti, Matteo
Brobston, Catherine
Hess, John
Leppert, Greg
Watson, Amanda
Zittrain, Jonathan
contents Publicly launched in 2004, the Google Books project has scanned tens of millions of items in partnership with libraries around the world. As part of this project, Google created the Google Return Interface (GRIN). Through this platform, libraries can access their scanned collections, the associated metadata, and the ongoing OCR and metadata improvements that become available as Google reprocesses these collections using new technologies. When downloading the Harvard Library Google Books collection from GRIN to develop the Institutional Books dataset, we encountered several challenges related to rate-limiting and atomized metadata within the GRIN platform. To overcome these challenges and help other libraries make more robust use of their Google Books collections, this technical report introduces the initial release of GRIN Transfer. This open-source and production-ready Python pipeline allows partner libraries to efficiently retrieve their Google Books collections from GRIN. This report also introduces an updated version of our Institutional Books 1.0 pipeline, initially used to analyze, augment, and assemble the Institutional Books 1.0 dataset. We have revised this pipeline for compatibility with the output format of GRIN Transfer. A library could pair these two tools to create an end-to-end processing pipeline for their Google Books collection to retrieve, structure, and enhance data available from GRIN. This report gives an overview of how GRIN Transfer was designed to optimize for reliability and usability in different environments, as well as guidance on configuration for various use cases.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11447
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GRIN Transfer: A production-ready tool for libraries to retrieve digital copies from Google Books
Daly, Liza
Cargnelutti, Matteo
Brobston, Catherine
Hess, John
Leppert, Greg
Watson, Amanda
Zittrain, Jonathan
Digital Libraries
Information Retrieval
Publicly launched in 2004, the Google Books project has scanned tens of millions of items in partnership with libraries around the world. As part of this project, Google created the Google Return Interface (GRIN). Through this platform, libraries can access their scanned collections, the associated metadata, and the ongoing OCR and metadata improvements that become available as Google reprocesses these collections using new technologies. When downloading the Harvard Library Google Books collection from GRIN to develop the Institutional Books dataset, we encountered several challenges related to rate-limiting and atomized metadata within the GRIN platform. To overcome these challenges and help other libraries make more robust use of their Google Books collections, this technical report introduces the initial release of GRIN Transfer. This open-source and production-ready Python pipeline allows partner libraries to efficiently retrieve their Google Books collections from GRIN. This report also introduces an updated version of our Institutional Books 1.0 pipeline, initially used to analyze, augment, and assemble the Institutional Books 1.0 dataset. We have revised this pipeline for compatibility with the output format of GRIN Transfer. A library could pair these two tools to create an end-to-end processing pipeline for their Google Books collection to retrieve, structure, and enhance data available from GRIN. This report gives an overview of how GRIN Transfer was designed to optimize for reliability and usability in different environments, as well as guidance on configuration for various use cases.
title GRIN Transfer: A production-ready tool for libraries to retrieve digital copies from Google Books
topic Digital Libraries
Information Retrieval
url https://arxiv.org/abs/2511.11447