TensorSocket: Shared Data Loading for Deep Learning Training
Fuente:
arXiv
Guardado en:
| Autores principales: | Robroek, Ties, Nielsen, Neil Kim, Tözün, Pınar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Modyn: Data-Centric Machine Learning Pipeline Orchestration
por: Böther, Maximilian, et al.
Publicado: (2023)
por: Böther, Maximilian, et al.
Publicado: (2023)
CARMA: Collocation-Aware Resource Manager
por: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Publicado: (2025)
por: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Publicado: (2025)
GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations
por: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Publicado: (2026)
por: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Publicado: (2026)
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training
por: Zhao, Pinxue, et al.
Publicado: (2024)
por: Zhao, Pinxue, et al.
Publicado: (2024)
10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
por: Afroz, Sabiha, et al.
Publicado: (2025)
por: Afroz, Sabiha, et al.
Publicado: (2025)
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
por: Li, Shigang, et al.
Publicado: (2019)
por: Li, Shigang, et al.
Publicado: (2019)
SparDL: Distributed Deep Learning Training with Efficient Sparse Communication
por: Zhao, Minjun, et al.
Publicado: (2023)
por: Zhao, Minjun, et al.
Publicado: (2023)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
por: Arfeen, Daiyaan, et al.
Publicado: (2025)
por: Arfeen, Daiyaan, et al.
Publicado: (2025)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
por: Woo, Sunghyeon, et al.
Publicado: (2026)
por: Woo, Sunghyeon, et al.
Publicado: (2026)
GraNNDis: Efficient Unified Distributed Training Framework for Deep GNNs on Large Clusters
por: Song, Jaeyong, et al.
Publicado: (2023)
por: Song, Jaeyong, et al.
Publicado: (2023)
Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
por: Dimlioglu, Tolga, et al.
Publicado: (2025)
por: Dimlioglu, Tolga, et al.
Publicado: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
por: Yarlagadda, Srihas, et al.
Publicado: (2025)
por: Yarlagadda, Srihas, et al.
Publicado: (2025)
ShardTensor: Domain Parallelism for Scientific Machine Learning
por: Adams, Corey, et al.
Publicado: (2026)
por: Adams, Corey, et al.
Publicado: (2026)
Two-dimensional Sparse Parallelism for Large Scale Deep Learning Recommendation Model Training
por: Zhang, Xin, et al.
Publicado: (2025)
por: Zhang, Xin, et al.
Publicado: (2025)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
por: Liu, Xinyi, et al.
Publicado: (2026)
por: Liu, Xinyi, et al.
Publicado: (2026)
Decoupled Vertical Federated Learning for Practical Training on Vertically Partitioned Data
por: Amalanshu, Avi, et al.
Publicado: (2024)
por: Amalanshu, Avi, et al.
Publicado: (2024)
Accelerating Communication in Deep Learning Recommendation Model Training with Dual-Level Adaptive Lossy Compression
por: Feng, Hao, et al.
Publicado: (2024)
por: Feng, Hao, et al.
Publicado: (2024)
MinatoLoader: Accelerating Machine Learning Training Through Efficient Data Preprocessing
por: Nouaji, Rahma, et al.
Publicado: (2025)
por: Nouaji, Rahma, et al.
Publicado: (2025)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
por: Xu, Jiale, et al.
Publicado: (2024)
por: Xu, Jiale, et al.
Publicado: (2024)
HeteroSwitch: Characterizing and Taming System-Induced Data Heterogeneity in Federated Learning
por: Kim, Gyudong, et al.
Publicado: (2024)
por: Kim, Gyudong, et al.
Publicado: (2024)
FedQueue: Queue-Aware Federated Learning for Cross-Facility HPC Training
por: Li, Yijiang, et al.
Publicado: (2026)
por: Li, Yijiang, et al.
Publicado: (2026)
Robust Federated Learning Mitigates Client-side Training Data Distribution Inference Attacks
por: Xu, Yichang, et al.
Publicado: (2024)
por: Xu, Yichang, et al.
Publicado: (2024)
Understanding Silent Data Corruption in LLM Training
por: Ma, Jeffrey, et al.
Publicado: (2025)
por: Ma, Jeffrey, et al.
Publicado: (2025)
BLOSSOM: Block-wise Federated Learning Over Shared and Sparse Observed Modalities
por: R, Pranav M, et al.
Publicado: (2026)
por: R, Pranav M, et al.
Publicado: (2026)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
por: Kim, Heehoon, et al.
Publicado: (2026)
por: Kim, Heehoon, et al.
Publicado: (2026)
Empowering Distributed Training with Sparsity-driven Data Synchronization
por: Wang, Zhuang, et al.
Publicado: (2023)
por: Wang, Zhuang, et al.
Publicado: (2023)
Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor Collections
por: Wagenländer, Marcel, et al.
Publicado: (2023)
por: Wagenländer, Marcel, et al.
Publicado: (2023)
Fused3S: Fast Sparse Attention on Tensor Cores
por: Li, Zitong, et al.
Publicado: (2025)
por: Li, Zitong, et al.
Publicado: (2025)
BLoad: Enhancing Neural Network Training with Efficient Sequential Data Handling
por: Ruschel, Raphael, et al.
Publicado: (2023)
por: Ruschel, Raphael, et al.
Publicado: (2023)
Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective
por: Yuan, Hao, et al.
Publicado: (2023)
por: Yuan, Hao, et al.
Publicado: (2023)
Training Machine Learning models at the Edge: A Survey
por: Khouas, Aymen Rayane, et al.
Publicado: (2024)
por: Khouas, Aymen Rayane, et al.
Publicado: (2024)
Energy-Aware Decentralized Learning with Intermittent Model Training
por: Dhasade, Akash, et al.
Publicado: (2024)
por: Dhasade, Akash, et al.
Publicado: (2024)
ReInc: Scaling Training of Dynamic Graph Neural Networks
por: Guan, Mingyu, et al.
Publicado: (2025)
por: Guan, Mingyu, et al.
Publicado: (2025)
Scaling State-Space Models on Multiple GPUs with Tensor Parallelism
por: Dutt, Anurag, et al.
Publicado: (2026)
por: Dutt, Anurag, et al.
Publicado: (2026)
Traversal Learning: A Lossless And Efficient Distributed Learning Framework
por: Batbaatar, Erdenebileg, et al.
Publicado: (2025)
por: Batbaatar, Erdenebileg, et al.
Publicado: (2025)
Revisiting Early-Learning Regularization When Federated Learning Meets Noisy Labels
por: Kim, Taehyeon, et al.
Publicado: (2024)
por: Kim, Taehyeon, et al.
Publicado: (2024)
Going Forward-Forward in Distributed Deep Learning
por: Aktemur, Ege, et al.
Publicado: (2024)
por: Aktemur, Ege, et al.
Publicado: (2024)
Aryl: An Elastic Cluster Scheduler for Deep Learning
por: Li, Jiamin, et al.
Publicado: (2022)
por: Li, Jiamin, et al.
Publicado: (2022)
FedAuxHMTL: Federated Auxiliary Hard-Parameter Sharing Multi-Task Learning for Network Edge Traffic Classification
por: Ahmed, Faisal, et al.
Publicado: (2024)
por: Ahmed, Faisal, et al.
Publicado: (2024)
Initialization Matters: Unraveling the Impact of Pre-Training on Federated Learning
por: Jhunjhunwala, Divyansh, et al.
Publicado: (2025)
por: Jhunjhunwala, Divyansh, et al.
Publicado: (2025)
Ejemplares similares
-
Modyn: Data-Centric Machine Learning Pipeline Orchestration
por: Böther, Maximilian, et al.
Publicado: (2023) -
CARMA: Collocation-Aware Resource Manager
por: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Publicado: (2025) -
GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations
por: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Publicado: (2026) -
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training
por: Zhao, Pinxue, et al.
Publicado: (2024) -
10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
por: Afroz, Sabiha, et al.
Publicado: (2025)