Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Desai, Omkar, Jiao, Ziyang, Pei, Shuyi, Bhimani, Janki, Kim, Bryan S.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918206555291648
author Desai, Omkar
Jiao, Ziyang
Pei, Shuyi
Bhimani, Janki
Kim, Bryan S.
author_facet Desai, Omkar
Jiao, Ziyang
Pei, Shuyi
Bhimani, Janki
Kim, Bryan S.
contents Input data preprocessing is a common bottleneck when concurrently training multimedia machine learning (ML) models in modern systems. To alleviate these bottlenecks and reduce the training time for concurrent jobs, we present Seneca, a data loading system that optimizes cache partitioning and data sampling for the data storage and ingestion (DSI) pipeline. The design of Seneca contains two key techniques. First, Seneca uses a performance model for the data pipeline to optimally partition the cache for three different forms of data (encoded, decoded, and augmented). Second, Seneca opportunistically serves cached data over uncached ones during random batch sampling so that concurrent jobs benefit from each other. We implement Seneca by modifying PyTorch and demonstrate its effectiveness by comparing it against several state-of-the-art caching systems for DNN training. Seneca reduces the makespan by 45.23% compared to PyTorch and increases data processing throughput by up to 3.45x compared to the next best dataloader.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13724
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
Desai, Omkar
Jiao, Ziyang
Pei, Shuyi
Bhimani, Janki
Kim, Bryan S.
Operating Systems
Artificial Intelligence
Machine Learning
Input data preprocessing is a common bottleneck when concurrently training multimedia machine learning (ML) models in modern systems. To alleviate these bottlenecks and reduce the training time for concurrent jobs, we present Seneca, a data loading system that optimizes cache partitioning and data sampling for the data storage and ingestion (DSI) pipeline. The design of Seneca contains two key techniques. First, Seneca uses a performance model for the data pipeline to optimally partition the cache for three different forms of data (encoded, decoded, and augmented). Second, Seneca opportunistically serves cached data over uncached ones during random batch sampling so that concurrent jobs benefit from each other. We implement Seneca by modifying PyTorch and demonstrate its effectiveness by comparing it against several state-of-the-art caching systems for DNN training. Seneca reduces the makespan by 45.23% compared to PyTorch and increases data processing throughput by up to 3.45x compared to the next best dataloader.
title Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
topic Operating Systems
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.13724