Declarative Data Pipeline for Large Scale ML Services
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911250842124288 |
|---|---|
| author | Yang, Yunzhao Wang, Runhui Liu, Xuanqing Krishnan, Adit Tao, Yefan Deng, Yuqian Yao, Kuangyou Sun, Peiyuan Johnson, Henrik sinha, Aditi Golac, Davor Friedland, Gerald Shakeel, Usman Cooke, Daryl Sullivan, Joe Chandrasekaran, Madhusudhanan Kong, Chris |
| author_facet | Yang, Yunzhao Wang, Runhui Liu, Xuanqing Krishnan, Adit Tao, Yefan Deng, Yuqian Yao, Kuangyou Sun, Peiyuan Johnson, Henrik sinha, Aditi Golac, Davor Friedland, Gerald Shakeel, Usman Cooke, Daryl Sullivan, Joe Chandrasekaran, Madhusudhanan Kong, Chris |
| contents | Modern distributed data processing systems struggle to balance performance, maintainability, and developer productivity when integrating machine learning at scale. These challenges intensify in large collaborative environments due to high communication overhead and coordination complexity. We present a "Declarative Data Pipeline" (DDP) architecture that addresses these challenges while processing billions of records efficiently. Our modular framework seamlessly integrates machine learning within Apache Spark using logical computation units called Pipes, departing from traditional microservice approaches. By establishing clear component boundaries and standardized interfaces, we achieve modularity and optimization without sacrificing maintainability. Enterprise case studies demonstrate substantial improvements: 50% better development efficiency, collaboration efforts compressed from weeks to days, 500x scalability improvement, and 10x throughput gains. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_15105 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Declarative Data Pipeline for Large Scale ML Services Yang, Yunzhao Wang, Runhui Liu, Xuanqing Krishnan, Adit Tao, Yefan Deng, Yuqian Yao, Kuangyou Sun, Peiyuan Johnson, Henrik sinha, Aditi Golac, Davor Friedland, Gerald Shakeel, Usman Cooke, Daryl Sullivan, Joe Chandrasekaran, Madhusudhanan Kong, Chris Distributed, Parallel, and Cluster Computing Modern distributed data processing systems struggle to balance performance, maintainability, and developer productivity when integrating machine learning at scale. These challenges intensify in large collaborative environments due to high communication overhead and coordination complexity. We present a "Declarative Data Pipeline" (DDP) architecture that addresses these challenges while processing billions of records efficiently. Our modular framework seamlessly integrates machine learning within Apache Spark using logical computation units called Pipes, departing from traditional microservice approaches. By establishing clear component boundaries and standardized interfaces, we achieve modularity and optimization without sacrificing maintainability. Enterprise case studies demonstrate substantial improvements: 50% better development efficiency, collaboration efforts compressed from weeks to days, 500x scalability improvement, and 10x throughput gains. |
| title | Declarative Data Pipeline for Large Scale ML Services |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2508.15105 |