Enabling Large Batch Size Training for DNN Models Beyond the Memory Limit While Maintaining Performance
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Piao, XinYu, Synn, DoangJoo, Park, JooYoung, Kim, Jong-Kook |
|---|---|
| Format: | Preprint |
| Publié: |
2021
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
par: Park, JooYoung, et autres
Publié: (2026)
par: Park, JooYoung, et autres
Publié: (2026)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024)
par: Chen, Jiabin, et autres
Publié: (2024)
Training DNN Models over Heterogeneous Clusters with Optimal Performance
par: Nie, Chengyi, et autres
Publié: (2024)
par: Nie, Chengyi, et autres
Publié: (2024)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
par: K., Prashanthi S., et autres
Publié: (2023)
par: K., Prashanthi S., et autres
Publié: (2023)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
par: Guo, Cong, et autres
Publié: (2024)
par: Guo, Cong, et autres
Publié: (2024)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
par: Lee, Munkyu, et autres
Publié: (2024)
par: Lee, Munkyu, et autres
Publié: (2024)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
par: Sakip, Akhmed, et autres
Publié: (2026)
par: Sakip, Akhmed, et autres
Publié: (2026)
Memory Efficient and Staleness Free Pipeline Parallel DNN Training Framework with Improved Convergence Speed
par: Dutta, Ankita, et autres
Publié: (2025)
par: Dutta, Ankita, et autres
Publié: (2025)
TiMePReSt: Time and Memory Efficient Pipeline Parallel DNN Training with Removed Staleness
par: Dutta, Ankita, et autres
Publié: (2024)
par: Dutta, Ankita, et autres
Publié: (2024)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
par: K., Prashanthi S., et autres
Publié: (2025)
par: K., Prashanthi S., et autres
Publié: (2025)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
par: Duan, Jiangfei, et autres
Publié: (2024)
par: Duan, Jiangfei, et autres
Publié: (2024)
SWIFT: Expedited Failure Recovery for Large-scale DNN Training
par: Zhong, Yuchen, et autres
Publié: (2023)
par: Zhong, Yuchen, et autres
Publié: (2023)
Nezha: Breaking Multi-Rail Network Barriers for Distributed DNN Training
par: Yu, Enda, et autres
Publié: (2024)
par: Yu, Enda, et autres
Publié: (2024)
A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
par: Jiang, Lijuan, et autres
Publié: (2025)
par: Jiang, Lijuan, et autres
Publié: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
par: Zhang, WenZheng, et autres
Publié: (2024)
par: Zhang, WenZheng, et autres
Publié: (2024)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
par: Chen, Lequn, et autres
Publié: (2023)
par: Chen, Lequn, et autres
Publié: (2023)
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
par: Zhang, Shiwei, et autres
Publié: (2024)
par: Zhang, Shiwei, et autres
Publié: (2024)
GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations
par: Yousefzadeh-Asl-Miandoab, Ehsan, et autres
Publié: (2026)
par: Yousefzadeh-Asl-Miandoab, Ehsan, et autres
Publié: (2026)
Maxing Out the SVM: Performance Impact of Memory and Program Cache Sizes in the Agave Validator
par: Vural, Turan, et autres
Publié: (2025)
par: Vural, Turan, et autres
Publié: (2025)
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
par: Lu, Zhengxian, et autres
Publié: (2024)
par: Lu, Zhengxian, et autres
Publié: (2024)
EaCO: Resource Sharing Dynamics and Its Impact on Energy Efficiency for DNN Training
par: Haghshenas, Kawsar, et autres
Publié: (2024)
par: Haghshenas, Kawsar, et autres
Publié: (2024)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
par: Svedas, Jonas, et autres
Publié: (2025)
par: Svedas, Jonas, et autres
Publié: (2025)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
par: Pang, Bowen, et autres
Publié: (2025)
par: Pang, Bowen, et autres
Publié: (2025)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
par: Ekelund, Jonah, et autres
Publié: (2025)
par: Ekelund, Jonah, et autres
Publié: (2025)
CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference
par: Zou, Yulin, et autres
Publié: (2026)
par: Zou, Yulin, et autres
Publié: (2026)
Collaborative Batch Size Optimization for Federated Learning
par: Geimer, Arno, et autres
Publié: (2025)
par: Geimer, Arno, et autres
Publié: (2025)
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
par: Lee, Haeun, et autres
Publié: (2025)
par: Lee, Haeun, et autres
Publié: (2025)
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
par: Song, Jaeyong, et autres
Publié: (2026)
par: Song, Jaeyong, et autres
Publié: (2026)
PaSE: Parallelization Strategies for Efficient DNN Training
par: Elango, Venmugil
Publié: (2024)
par: Elango, Venmugil
Publié: (2024)
Practical Performance Guarantees for Pipelined DNN Inference
par: Archer, Aaron, et autres
Publié: (2023)
par: Archer, Aaron, et autres
Publié: (2023)
On Optimal Batch Size in Coded Computing
par: Saha, Swapnil, et autres
Publié: (2025)
par: Saha, Swapnil, et autres
Publié: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
par: Cao, Jiahe, et autres
Publié: (2026)
par: Cao, Jiahe, et autres
Publié: (2026)
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
par: Jeon, Byungsoo, et autres
Publié: (2024)
par: Jeon, Byungsoo, et autres
Publié: (2024)
Graph Neural Network Training Systems: A Performance Comparison of Full-Graph and Mini-Batch
par: Bajaj, Saurabh, et autres
Publié: (2024)
par: Bajaj, Saurabh, et autres
Publié: (2024)
Pagoda: An Energy and Time Roofline Study for DNN Workloads on Edge Accelerators
par: K., Prashanthi S., et autres
Publié: (2025)
par: K., Prashanthi S., et autres
Publié: (2025)
Collaborative Satellite Computing through Adaptive DNN Task Splitting and Offloading
par: Peng, Shifeng, et autres
Publié: (2024)
par: Peng, Shifeng, et autres
Publié: (2024)
Ocularone-Bench: Benchmarking DNN Models on GPUs to Assist the Visually Impaired
par: Raj, Suman, et autres
Publié: (2025)
par: Raj, Suman, et autres
Publié: (2025)
EcoFed: Efficient Communication for DNN Partitioning-based Federated Learning
par: Wu, Di, et autres
Publié: (2023)
par: Wu, Di, et autres
Publié: (2023)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
par: Li, Rui, et autres
Publié: (2024)
par: Li, Rui, et autres
Publié: (2024)
Documents similaires
-
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
par: Park, JooYoung, et autres
Publié: (2026) -
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024) -
Training DNN Models over Heterogeneous Clusters with Optimal Performance
par: Nie, Chengyi, et autres
Publié: (2024) -
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
par: K., Prashanthi S., et autres
Publié: (2023) -
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
par: Guo, Cong, et autres
Publié: (2024)