Provenance Tracking in Large-Scale Machine Learning Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Padovani, Gabriele, Anantharaj, Valentine, Fiore, Sandro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
yProv4ML: Effortless Provenance Tracking for Machine Learning Systems
by: Padovani, Gabriele, et al.
Published: (2025)
by: Padovani, Gabriele, et al.
Published: (2025)
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
by: Kokolis, Apostolos, et al.
Published: (2024)
by: Kokolis, Apostolos, et al.
Published: (2024)
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
by: Wang, Weixun, et al.
Published: (2025)
by: Wang, Weixun, et al.
Published: (2025)
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
by: Chai, Huichao, et al.
Published: (2026)
by: Chai, Huichao, et al.
Published: (2026)
Leveraging Neural Graph Compilers in Machine Learning Research for Edge-Cloud Systems
by: Furutanpey, Alireza, et al.
Published: (2025)
by: Furutanpey, Alireza, et al.
Published: (2025)
Two-dimensional Sparse Parallelism for Large Scale Deep Learning Recommendation Model Training
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Federated Koopman-Reservoir Learning for Large-Scale Multivariate Time-Series Anomaly Detection
by: Le, Long Tan, et al.
Published: (2025)
by: Le, Long Tan, et al.
Published: (2025)
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
Personalized Semi-Supervised Federated Learning for Human Activity Recognition
by: Presotto, Riccardo, et al.
Published: (2021)
by: Presotto, Riccardo, et al.
Published: (2021)
DYNAMIX: RL-based Adaptive Batch Size Optimization in Distributed Machine Learning Systems
by: Dai, Yuanjun, et al.
Published: (2025)
by: Dai, Yuanjun, et al.
Published: (2025)
ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning
by: Song, Jingwei, et al.
Published: (2026)
by: Song, Jingwei, et al.
Published: (2026)
Robust Decentralized Learning with Local Updates and Gradient Tracking
by: Ghiasvand, Sajjad, et al.
Published: (2024)
by: Ghiasvand, Sajjad, et al.
Published: (2024)
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
by: Lian, Xinyu, et al.
Published: (2024)
by: Lian, Xinyu, et al.
Published: (2024)
Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core Systems
by: Xia, Yufan, et al.
Published: (2024)
by: Xia, Yufan, et al.
Published: (2024)
EARL: Efficient Agentic Reinforcement Learning Systems for Large Language Models
by: Tan, Zheyue, et al.
Published: (2025)
by: Tan, Zheyue, et al.
Published: (2025)
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
by: Jiang, Ziheng, et al.
Published: (2024)
by: Jiang, Ziheng, et al.
Published: (2024)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
by: Jin, Chao, et al.
Published: (2025)
by: Jin, Chao, et al.
Published: (2025)
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
by: Hagemann, Johannes, et al.
Published: (2023)
by: Hagemann, Johannes, et al.
Published: (2023)
Minder: Faulty Machine Detection for Large-scale Distributed Model Training
by: Deng, Yangtao, et al.
Published: (2024)
by: Deng, Yangtao, et al.
Published: (2024)
Machine Learning for Consistency Violation Faults Analysis
by: Giri, Kamal, et al.
Published: (2025)
by: Giri, Kamal, et al.
Published: (2025)
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025)
by: Lian, Xinyu, et al.
Published: (2025)
ShardTensor: Domain Parallelism for Scientific Machine Learning
by: Adams, Corey, et al.
Published: (2026)
by: Adams, Corey, et al.
Published: (2026)
Heterogeneity: An Open Challenge for Federated On-board Machine Learning
by: Hartmann, Maria, et al.
Published: (2024)
by: Hartmann, Maria, et al.
Published: (2024)
Training Machine Learning models at the Edge: A Survey
by: Khouas, Aymen Rayane, et al.
Published: (2024)
by: Khouas, Aymen Rayane, et al.
Published: (2024)
PiPar: Pipeline Parallelism for Collaborative Machine Learning
by: Zhang, Zihan, et al.
Published: (2022)
by: Zhang, Zihan, et al.
Published: (2022)
Algorithms for Collaborative Machine Learning under Statistical Heterogeneity
by: Hahn, Seok-Ju
Published: (2024)
by: Hahn, Seok-Ju
Published: (2024)
Large-Scale Graph Building in Dynamic Environments: Low Latency and High Quality
by: de Almeida, Filipe Miguel Gonçalves, et al.
Published: (2025)
by: de Almeida, Filipe Miguel Gonçalves, et al.
Published: (2025)
Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
by: Waleffe, Roger, et al.
Published: (2025)
by: Waleffe, Roger, et al.
Published: (2025)
A Semantic Partitioning Method for Large-Scale Training of Knowledge Graph Embeddings
by: Bai, Yuhe
Published: (2025)
by: Bai, Yuhe
Published: (2025)
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
by: Chen, Ling, et al.
Published: (2026)
by: Chen, Ling, et al.
Published: (2026)
Federated Behavioural Planes: Explaining the Evolution of Client Behaviour in Federated Learning
by: Fenoglio, Dario, et al.
Published: (2024)
by: Fenoglio, Dario, et al.
Published: (2024)
GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism
by: Polisetty, Sandeep, et al.
Published: (2023)
by: Polisetty, Sandeep, et al.
Published: (2023)
Unlearning during Learning: An Efficient Federated Machine Unlearning Method
by: Gu, Hanlin, et al.
Published: (2024)
by: Gu, Hanlin, et al.
Published: (2024)
Machine Learning-Based Research on the Adaptability of Adolescents to Online Education
by: Wang, Mingwei, et al.
Published: (2024)
by: Wang, Mingwei, et al.
Published: (2024)
A Robust Federated Learning Framework for Undependable Devices at Scale
by: Wang, Shilong, et al.
Published: (2024)
by: Wang, Shilong, et al.
Published: (2024)
Dependency Aware Incident Linking in Large Cloud Systems
by: Ghosh, Supriyo, et al.
Published: (2024)
by: Ghosh, Supriyo, et al.
Published: (2024)
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
by: Jiang, Zhida, et al.
Published: (2026)
by: Jiang, Zhida, et al.
Published: (2026)
A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication
by: Xia, Yufan, et al.
Published: (2026)
by: Xia, Yufan, et al.
Published: (2026)
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
Similar Items
-
yProv4ML: Effortless Provenance Tracking for Machine Learning Systems
by: Padovani, Gabriele, et al.
Published: (2025) -
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
by: Kokolis, Apostolos, et al.
Published: (2024) -
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
by: Wang, Weixun, et al.
Published: (2025) -
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
by: Chai, Huichao, et al.
Published: (2026) -
Leveraging Neural Graph Compilers in Machine Learning Research for Edge-Cloud Systems
by: Furutanpey, Alireza, et al.
Published: (2025)