MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
Fuente:
arXiv
Saved in:
| Main Authors: | Maurya, Avinash, Rafique, M. Mustafa, Cappello, Franck, Nicolae, Bogdan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
by: Maurya, Avinash, et al.
Published: (2026)
by: Maurya, Avinash, et al.
Published: (2026)
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
Offloading tracing for real-time systems using a scalable cloud infrastructure
by: Schmidt, David Jannis, et al.
Published: (2025)
by: Schmidt, David Jannis, et al.
Published: (2025)
CodeCRDT: Observation-Driven Coordination for Multi-Agent LLM Code Generation
by: Pugachev, Sergey
Published: (2025)
by: Pugachev, Sergey
Published: (2025)
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
by: Xu, Youxuan, et al.
Published: (2025)
by: Xu, Youxuan, et al.
Published: (2025)
Combinatorial Client-Master Multiagent Deep Reinforcement Learning for Task Offloading in Mobile Edge Computing
by: Gebrekidan, Tesfay Zemuy, et al.
Published: (2024)
by: Gebrekidan, Tesfay Zemuy, et al.
Published: (2024)
Asynchronous Multi-Server Federated Learning for Geo-Distributed Clients
by: Zuo, Yuncong, et al.
Published: (2024)
by: Zuo, Yuncong, et al.
Published: (2024)
Towards Optimal Heterogeneous Client Sampling in Multi-Model Federated Learning
by: Zhang, Haoran, et al.
Published: (2025)
by: Zhang, Haoran, et al.
Published: (2025)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
by: Bian, Zhuohang, et al.
Published: (2025)
by: Bian, Zhuohang, et al.
Published: (2025)
Adaptive GPU Resource Allocation for Multi-Agent Collaborative Reasoning in Serverless Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
by: Murimi, Almond Kiruthu
Published: (2025)
by: Murimi, Almond Kiruthu
Published: (2025)
Separating Intelligence from Execution: A Workflow Engine for the Model Context Protocol
by: Parmar, Abhinav Singh
Published: (2026)
by: Parmar, Abhinav Singh
Published: (2026)
Aergia: Leveraging Heterogeneity in Federated Learning Systems
by: Cox, Bart, et al.
Published: (2022)
by: Cox, Bart, et al.
Published: (2022)
Roadmap for Edge AI: A Dagstuhl Perspective
by: Ding, Aaron Yi, et al.
Published: (2021)
by: Ding, Aaron Yi, et al.
Published: (2021)
Parameterizing Federated Continual Learning for Reproducible Research
by: Cox, Bart, et al.
Published: (2024)
by: Cox, Bart, et al.
Published: (2024)
Asynchronous Byzantine Federated Learning
by: Cox, Bart, et al.
Published: (2024)
by: Cox, Bart, et al.
Published: (2024)
Hyper-parameter Optimization for Federated Learning with Step-wise Adaptive Mechanism
by: Saadati, Yasaman, et al.
Published: (2024)
by: Saadati, Yasaman, et al.
Published: (2024)
A Domain-Driven Design Simulator for Business Logic-Rich Microservice Systems
by: Pereira, Daniel da Palma, et al.
Published: (2026)
by: Pereira, Daniel da Palma, et al.
Published: (2026)
Training Diffusion Models with Federated Learning
by: de Goede, Matthijs, et al.
Published: (2024)
by: de Goede, Matthijs, et al.
Published: (2024)
Connecting Large Language Model Agent to High Performance Computing Resource
by: Ma, Heng, et al.
Published: (2025)
by: Ma, Heng, et al.
Published: (2025)
Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees
by: Xin, Jihao, et al.
Published: (2023)
by: Xin, Jihao, et al.
Published: (2023)
Streaming REST APIs for Large Financial Transaction Exports from Relational Databases
by: Kandiraju, Abhiram
Published: (2026)
by: Kandiraju, Abhiram
Published: (2026)
Comparison of Autoscaling Frameworks for Containerised Machine-Learning-Applications in a Local and Cloud Environment
by: Schroeder, Christian, et al.
Published: (2023)
by: Schroeder, Christian, et al.
Published: (2023)
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
by: Kang, Daemyung, et al.
Published: (2026)
by: Kang, Daemyung, et al.
Published: (2026)
A Full Compression Pipeline for Green Federated Learning in Communication-Constrained Environments
by: Colybes, Elouan, et al.
Published: (2026)
by: Colybes, Elouan, et al.
Published: (2026)
Uncertainty Estimation in Multi-Agent Distributed Learning for AI-Enabled Edge Devices
by: Radchenko, Gleb, et al.
Published: (2024)
by: Radchenko, Gleb, et al.
Published: (2024)
Learning In Chaos: Efficient Autoscaling and Self-Healing for Multi-Party Distributed Training
by: Feng, Wenjiao, et al.
Published: (2025)
by: Feng, Wenjiao, et al.
Published: (2025)
GPU-Augmented OLAP Execution Engine: GPU Offloading
by: Chang, Ilsun
Published: (2025)
by: Chang, Ilsun
Published: (2025)
Agentic Compilation: Mitigating the LLM Rerun Crisis for Minimized-Inference-Cost Web Automation
by: Chundru, Jagadeesh
Published: (2026)
by: Chundru, Jagadeesh
Published: (2026)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
by: Wang, Zheng, et al.
Published: (2025)
by: Wang, Zheng, et al.
Published: (2025)
TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
Complex Event Processing in the Edge: A Combined Optimization Approach for Data and Code Placement
by: Uyanık, Halit, et al.
Published: (2026)
by: Uyanık, Halit, et al.
Published: (2026)
A Framework for the Interoperability of Cloud Platforms: Towards FAIR Data in SAFE Environments
by: Grossman, Robert L., et al.
Published: (2022)
by: Grossman, Robert L., et al.
Published: (2022)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
HFedATM: Hierarchical Federated Domain Generalization via Optimal Transport and Regularized Mean Aggregation
by: Nguyen, Thinh, et al.
Published: (2025)
by: Nguyen, Thinh, et al.
Published: (2025)
ADF-LoRA: Alternating Low-Rank Aggregation for Decentralized Federated Fine-Tuning
by: Wang, Xiaoyu, et al.
Published: (2025)
by: Wang, Xiaoyu, et al.
Published: (2025)
DAGER: Exact Gradient Inversion for Large Language Models
by: Petrov, Ivo, et al.
Published: (2024)
by: Petrov, Ivo, et al.
Published: (2024)
Similar Items
-
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
by: Maurya, Avinash, et al.
Published: (2024) -
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
by: Maurya, Avinash, et al.
Published: (2026) -
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
by: Maurya, Avinash, et al.
Published: (2024) -
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
by: Maurya, Avinash, et al.
Published: (2024) -
Offloading tracing for real-time systems using a scalable cloud infrastructure
by: Schmidt, David Jannis, et al.
Published: (2025)