A Comparative Analysis of Distributed Training Strategies for GPT-2
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Patwardhan, Ishan, Gandhi, Shubham, Khare, Om, Joshi, Amit, Sawant, Suraj |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hybrid Quantum-HPC Solutions for Max-Cut: Bridging Classical and Quantum Algorithms
von: Patwardhan, Ishan, et al.
Veröffentlicht: (2024)
von: Patwardhan, Ishan, et al.
Veröffentlicht: (2024)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
von: Patel, Ishan, et al.
Veröffentlicht: (2026)
von: Patel, Ishan, et al.
Veröffentlicht: (2026)
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
Distributed Locking: Performance Analysis and Optimization Strategies
von: Rodriguez, Andre, et al.
Veröffentlicht: (2025)
von: Rodriguez, Andre, et al.
Veröffentlicht: (2025)
Comparative Analysis of Distributed Caching Algorithms: Performance Metrics and Implementation Considerations
von: Mayer, Helen, et al.
Veröffentlicht: (2025)
von: Mayer, Helen, et al.
Veröffentlicht: (2025)
Comparative Analysis of Lightweight Kubernetes Distributions for Edge Computing: Performance and Resource Efficiency
von: Yakubov, Diyaz, et al.
Veröffentlicht: (2025)
von: Yakubov, Diyaz, et al.
Veröffentlicht: (2025)
Cost-Performance Analysis: A Comparative Study of CPU-Based Serverless and GPU-Based Training Architectures
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
von: Gandhi, Rohan, et al.
Veröffentlicht: (2024)
von: Gandhi, Rohan, et al.
Veröffentlicht: (2024)
ParEval-Repo: A Benchmark Suite for Evaluating LLMs with Repository-level HPC Translation Tasks
von: Davis, Joshua H., et al.
Veröffentlicht: (2025)
von: Davis, Joshua H., et al.
Veröffentlicht: (2025)
Efficient Distributed MLLM Training with Cornstarch
von: Jang, Insu, et al.
Veröffentlicht: (2025)
von: Jang, Insu, et al.
Veröffentlicht: (2025)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
A Study on the Performance of Distributed Training of Data-driven CFD Simulations
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
A flexible FPGA accelerator for convolutional neural networks
von: Majumder, Kingshuk, et al.
Veröffentlicht: (2019)
von: Majumder, Kingshuk, et al.
Veröffentlicht: (2019)
Heta: Distributed Training of Heterogeneous Graph Neural Networks
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
Galvatron: Automatic Distributed Training for Large Transformer Models
von: Gumaan, Esmail
Veröffentlicht: (2025)
von: Gumaan, Esmail
Veröffentlicht: (2025)
Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain
von: Jang, Insu, et al.
Veröffentlicht: (2026)
von: Jang, Insu, et al.
Veröffentlicht: (2026)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
Optimizing Distributed Training Approaches for Scaling Neural Networks
von: Baligodugula, Vishnu Vardhan, et al.
Veröffentlicht: (2025)
von: Baligodugula, Vishnu Vardhan, et al.
Veröffentlicht: (2025)
Fault-Tolerant Decentralized Distributed Asynchronous Federated Learning with Adaptive Termination Detection
von: Akkinepally, Phani Sahasra, et al.
Veröffentlicht: (2025)
von: Akkinepally, Phani Sahasra, et al.
Veröffentlicht: (2025)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
Characterizing FaaS Workflows on Public Clouds: The Good, the Bad and the Ugly
von: Kulkarni, Varad, et al.
Veröffentlicht: (2025)
von: Kulkarni, Varad, et al.
Veröffentlicht: (2025)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
von: Shubham, et al.
Veröffentlicht: (2024)
von: Shubham, et al.
Veröffentlicht: (2024)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
von: Svedas, Jonas, et al.
Veröffentlicht: (2025)
von: Svedas, Jonas, et al.
Veröffentlicht: (2025)
FailSafe: High-performance Resilient Serving
von: Xu, Ziyi, et al.
Veröffentlicht: (2025)
von: Xu, Ziyi, et al.
Veröffentlicht: (2025)
Tetris: Efficient Intra-Datacenter Calls Packing for Large Conferencing Services
von: Gandhi, Rohan, et al.
Veröffentlicht: (2025)
von: Gandhi, Rohan, et al.
Veröffentlicht: (2025)
A Comparative Analysis of Identifier Schemes: UUIDv4, UUIDv7, and ULID for Distributed Systems
von: Kakolaki, Nima Karimian
Veröffentlicht: (2025)
von: Kakolaki, Nima Karimian
Veröffentlicht: (2025)
A Comparative Evaluation of Automated Analysis Tools for Solidity Smart Contracts
von: Wei, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Wei, Zhiyuan, et al.
Veröffentlicht: (2023)
Nezha: Breaking Multi-Rail Network Barriers for Distributed DNN Training
von: Yu, Enda, et al.
Veröffentlicht: (2024)
von: Yu, Enda, et al.
Veröffentlicht: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
von: Wang, Zhixin, et al.
Veröffentlicht: (2025)
von: Wang, Zhixin, et al.
Veröffentlicht: (2025)
PruneX: A Hierarchical Communication-Efficient System for Distributed CNN Training with Structured Pruning
von: Olama, Alireza, et al.
Veröffentlicht: (2025)
von: Olama, Alireza, et al.
Veröffentlicht: (2025)
FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
von: Gao, Yunqi, et al.
Veröffentlicht: (2025)
von: Gao, Yunqi, et al.
Veröffentlicht: (2025)
Analysis of Distributed Algorithms for Big-data
von: Purohit, Rajendra, et al.
Veröffentlicht: (2024)
von: Purohit, Rajendra, et al.
Veröffentlicht: (2024)
How to Evaluate Distributed Coordination Systems? -- A Survey and Analysis
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
MTGenRec: An Efficient Distributed Training System for Generative Recommendation Models in Meituan
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
von: Yao, Chenxuan, et al.
Veröffentlicht: (2025)
von: Yao, Chenxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hybrid Quantum-HPC Solutions for Max-Cut: Bridging Classical and Quantum Algorithms
von: Patwardhan, Ishan, et al.
Veröffentlicht: (2024) -
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
von: Patel, Ishan, et al.
Veröffentlicht: (2026) -
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024) -
Distributed Locking: Performance Analysis and Optimization Strategies
von: Rodriguez, Andre, et al.
Veröffentlicht: (2025) -
Comparative Analysis of Distributed Caching Algorithms: Performance Metrics and Implementation Considerations
von: Mayer, Helen, et al.
Veröffentlicht: (2025)