Data-Prep-Kit: getting your data ready for LLM application development

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wood, David, Lublinsky, Boris, Roytman, Alexy, Singh, Shivdeep, Adam, Constantin, Adebayo, Abdulhamid, An, Sungeun, Chang, Yuan Chi, Dang, Xuan-Hong, Desai, Nirmit, Dolfi, Michele, Emami-Gohari, Hajar, Eres, Revital, Goto, Takuya, Joshi, Dhiraj, Koyfman, Yan, Nassar, Mohammad, Patel, Hima, Selvam, Paramesvaran, Shah, Yousaf, Surendran, Saptha, Tsuzuku, Daiki, Zerfos, Petros, Daijavad, Shahrokh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929589130887168
author Wood, David
Lublinsky, Boris
Roytman, Alexy
Singh, Shivdeep
Adam, Constantin
Adebayo, Abdulhamid
An, Sungeun
Chang, Yuan Chi
Dang, Xuan-Hong
Desai, Nirmit
Dolfi, Michele
Emami-Gohari, Hajar
Eres, Revital
Goto, Takuya
Joshi, Dhiraj
Koyfman, Yan
Nassar, Mohammad
Patel, Hima
Selvam, Paramesvaran
Shah, Yousaf
Surendran, Saptha
Tsuzuku, Daiki
Zerfos, Petros
Daijavad, Shahrokh
author_facet Wood, David
Lublinsky, Boris
Roytman, Alexy
Singh, Shivdeep
Adam, Constantin
Adebayo, Abdulhamid
An, Sungeun
Chang, Yuan Chi
Dang, Xuan-Hong
Desai, Nirmit
Dolfi, Michele
Emami-Gohari, Hajar
Eres, Revital
Goto, Takuya
Joshi, Dhiraj
Koyfman, Yan
Nassar, Mohammad
Patel, Hima
Selvam, Paramesvaran
Shah, Yousaf
Surendran, Saptha
Tsuzuku, Daiki
Zerfos, Petros
Daijavad, Shahrokh
contents Data preparation is the first and a very important step towards any Large Language Model (LLM) development. This paper introduces an easy-to-use, extensible, and scale-flexible open-source data preparation toolkit called Data Prep Kit (DPK). DPK is architected and designed to enable users to scale their data preparation to their needs. With DPK they can prepare data on a local machine or effortlessly scale to run on a cluster with thousands of CPU Cores. DPK comes with a highly scalable, yet extensible set of modules that transform natural language and code data. If the user needs additional transforms, they can be easily developed using extensive DPK support for transform creation. These modules can be used independently or pipelined to perform a series of operations. In this paper, we describe DPK architecture and show its performance from a small scale to a very large number of CPUs. The modules from DPK have been used for the preparation of Granite Models [1] [2]. We believe DPK is a valuable contribution to the AI community to easily prepare data to enhance the performance of their LLM models or to fine-tune models with Retrieval-Augmented Generation (RAG).
format Preprint
id arxiv_https___arxiv_org_abs_2409_18164
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data-Prep-Kit: getting your data ready for LLM application development
Wood, David
Lublinsky, Boris
Roytman, Alexy
Singh, Shivdeep
Adam, Constantin
Adebayo, Abdulhamid
An, Sungeun
Chang, Yuan Chi
Dang, Xuan-Hong
Desai, Nirmit
Dolfi, Michele
Emami-Gohari, Hajar
Eres, Revital
Goto, Takuya
Joshi, Dhiraj
Koyfman, Yan
Nassar, Mohammad
Patel, Hima
Selvam, Paramesvaran
Shah, Yousaf
Surendran, Saptha
Tsuzuku, Daiki
Zerfos, Petros
Daijavad, Shahrokh
Artificial Intelligence
Computation and Language
Machine Learning
Data preparation is the first and a very important step towards any Large Language Model (LLM) development. This paper introduces an easy-to-use, extensible, and scale-flexible open-source data preparation toolkit called Data Prep Kit (DPK). DPK is architected and designed to enable users to scale their data preparation to their needs. With DPK they can prepare data on a local machine or effortlessly scale to run on a cluster with thousands of CPU Cores. DPK comes with a highly scalable, yet extensible set of modules that transform natural language and code data. If the user needs additional transforms, they can be easily developed using extensive DPK support for transform creation. These modules can be used independently or pipelined to perform a series of operations. In this paper, we describe DPK architecture and show its performance from a small scale to a very large number of CPUs. The modules from DPK have been used for the preparation of Granite Models [1] [2]. We believe DPK is a valuable contribution to the AI community to easily prepare data to enhance the performance of their LLM models or to fine-tune models with Retrieval-Augmented Generation (RAG).
title Data-Prep-Kit: getting your data ready for LLM application development
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2409.18164