ProTrain: Efficient LLM Training via Memory-Aware Techniques
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Hanmei, Zhou, Jin, Fu, Yao, Wang, Xiaoqun, Roane, Ramine, Guan, Hui, Liu, Tongping |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025)
di: Karfakis, George, et al.
Pubblicazione: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
di: Yuan, Ziqi, et al.
Pubblicazione: (2025)
di: Yuan, Ziqi, et al.
Pubblicazione: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
di: Lin, Mao, et al.
Pubblicazione: (2026)
di: Lin, Mao, et al.
Pubblicazione: (2026)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
di: Muhammad, Said, et al.
Pubblicazione: (2025)
di: Muhammad, Said, et al.
Pubblicazione: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
di: Sarkar, Aishwarya, et al.
Pubblicazione: (2024)
di: Sarkar, Aishwarya, et al.
Pubblicazione: (2024)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
di: Papavasileiou, Ioannis, et al.
Pubblicazione: (2026)
di: Papavasileiou, Ioannis, et al.
Pubblicazione: (2026)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
di: Zhang, Lingqi, et al.
Pubblicazione: (2025)
di: Zhang, Lingqi, et al.
Pubblicazione: (2025)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
di: Proaño, Andrès Rubio, et al.
Pubblicazione: (2024)
di: Proaño, Andrès Rubio, et al.
Pubblicazione: (2024)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
di: Shi, Jiabo, et al.
Pubblicazione: (2025)
di: Shi, Jiabo, et al.
Pubblicazione: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
di: Zhao, Yanbo, et al.
Pubblicazione: (2025)
di: Zhao, Yanbo, et al.
Pubblicazione: (2025)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
di: Lacey, Dane C., et al.
Pubblicazione: (2024)
di: Lacey, Dane C., et al.
Pubblicazione: (2024)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
di: Gupta, Ahan, et al.
Pubblicazione: (2026)
di: Gupta, Ahan, et al.
Pubblicazione: (2026)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
di: Ng, Nathan, et al.
Pubblicazione: (2026)
di: Ng, Nathan, et al.
Pubblicazione: (2026)
Energy-Aware Computing in the Year 2026
di: Tchakoute, Roblex Nana, et al.
Pubblicazione: (2026)
di: Tchakoute, Roblex Nana, et al.
Pubblicazione: (2026)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
di: Wahlgren, Jacob, et al.
Pubblicazione: (2025)
di: Wahlgren, Jacob, et al.
Pubblicazione: (2025)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
di: Sun, Tingyang, et al.
Pubblicazione: (2026)
di: Sun, Tingyang, et al.
Pubblicazione: (2026)
Communication-Aware Diffusion Load Balancing for Persistently Interacting Objects
di: Taylor, Maya, et al.
Pubblicazione: (2026)
di: Taylor, Maya, et al.
Pubblicazione: (2026)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
di: Fan, Ruibo, et al.
Pubblicazione: (2026)
di: Fan, Ruibo, et al.
Pubblicazione: (2026)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
di: Dutt, Anurag, et al.
Pubblicazione: (2025)
di: Dutt, Anurag, et al.
Pubblicazione: (2025)
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
di: Battaglini-Fischer, Sándor, et al.
Pubblicazione: (2025)
di: Battaglini-Fischer, Sándor, et al.
Pubblicazione: (2025)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
di: Qi, S., et al.
Pubblicazione: (2024)
di: Qi, S., et al.
Pubblicazione: (2024)
TrainMover: An Interruption-Resilient Runtime for ML Training
di: Lao, ChonLam, et al.
Pubblicazione: (2024)
di: Lao, ChonLam, et al.
Pubblicazione: (2024)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
di: Maurya, Avinash, et al.
Pubblicazione: (2026)
di: Maurya, Avinash, et al.
Pubblicazione: (2026)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
di: Boudaoud, Afif, et al.
Pubblicazione: (2026)
di: Boudaoud, Afif, et al.
Pubblicazione: (2026)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
di: Shan, Baodi, et al.
Pubblicazione: (2024)
di: Shan, Baodi, et al.
Pubblicazione: (2024)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
di: Liu, Shifang, et al.
Pubblicazione: (2025)
di: Liu, Shifang, et al.
Pubblicazione: (2025)
Efficient Serverless Cold Start: Reducing Library Loading Overhead by Profile-guided Optimization
di: Tariq, Syed Salauddin Mohammad, et al.
Pubblicazione: (2025)
di: Tariq, Syed Salauddin Mohammad, et al.
Pubblicazione: (2025)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
di: Mathews, Dhanya R, et al.
Pubblicazione: (2025)
di: Mathews, Dhanya R, et al.
Pubblicazione: (2025)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
di: Mao, Ying, et al.
Pubblicazione: (2020)
di: Mao, Ying, et al.
Pubblicazione: (2020)
Documenti analoghi
-
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025) -
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
di: Wang, Yuxin, et al.
Pubblicazione: (2023) -
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
di: Yuan, Ziqi, et al.
Pubblicazione: (2025) -
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
di: Lin, Mao, et al.
Pubblicazione: (2026) -
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
di: Muhammad, Said, et al.
Pubblicazione: (2025)