BladeDISC++: Memory Optimizations Based On Symbolic Shape
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yuan, Xiulong, Yan, Xu, Shen, Wenting, Qiu, Xiafei, Wang, Ang, Zhang, Jie, Li, Yong, Lin, Wei |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Efficient Long Context Fine-tuning with Chunk Flow
par: Yuan, Xiulong, et autres
Publié: (2025)
par: Yuan, Xiulong, et autres
Publié: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
par: Sun, Mingyu, et autres
Publié: (2025)
par: Sun, Mingyu, et autres
Publié: (2025)
MemAscend: System Memory Optimization for SSD-Offloaded LLM Fine-Tuning
par: Liaw, Yong-Cheng, et autres
Publié: (2025)
par: Liaw, Yong-Cheng, et autres
Publié: (2025)
Analysis and Optimized CXL-Attached Memory Allocation for Long-Context LLM Fine-Tuning
par: Liaw, Yong-Cheng, et autres
Publié: (2025)
par: Liaw, Yong-Cheng, et autres
Publié: (2025)
Accelerating Compound LLM Training Workloads with Maestro
par: Yuan, Xiulong, et autres
Publié: (2026)
par: Yuan, Xiulong, et autres
Publié: (2026)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
par: Lin, Bin, et autres
Publié: (2024)
par: Lin, Bin, et autres
Publié: (2024)
Bandwidth-Aware Network Topology Optimization for Decentralized Learning
par: Shen, Yipeng, et autres
Publié: (2025)
par: Shen, Yipeng, et autres
Publié: (2025)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
par: Maurya, Avinash, et autres
Publié: (2024)
par: Maurya, Avinash, et autres
Publié: (2024)
TierBase: A Workload-Driven Cost-Optimized Key-Value Store
par: Shen, Zhitao, et autres
Publié: (2025)
par: Shen, Zhitao, et autres
Publié: (2025)
SkyMemory: A LEO Edge Cache for Transformer Inference Optimization and Scale Out
par: Sandholm, Thomas, et autres
Publié: (2025)
par: Sandholm, Thomas, et autres
Publié: (2025)
PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping
par: Zhao, Zhixin, et autres
Publié: (2026)
par: Zhao, Zhixin, et autres
Publié: (2026)
SHARE: Optimizing Secure Hub Allocation and Routing Efficiency in Payment Channel Networks
par: Yang, Lingxiao, et autres
Publié: (2025)
par: Yang, Lingxiao, et autres
Publié: (2025)
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
par: Xu, Yi, et autres
Publié: (2024)
par: Xu, Yi, et autres
Publié: (2024)
Bridging Cache-Friendliness and Concurrency: A Locality-Optimized In-Memory B-Skiplist
par: Luo, Yicong, et autres
Publié: (2025)
par: Luo, Yicong, et autres
Publié: (2025)
Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
par: Zhao, Zhixin, et autres
Publié: (2024)
par: Zhao, Zhixin, et autres
Publié: (2024)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
par: Zhang, Yaozheng, et autres
Publié: (2025)
par: Zhang, Yaozheng, et autres
Publié: (2025)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
par: Ma, Chenxiang, et autres
Publié: (2025)
par: Ma, Chenxiang, et autres
Publié: (2025)
Optimizing Memory Allocation in Distributed Clusters with Predictive Modeling
par: Bader, Jonathan, et autres
Publié: (2026)
par: Bader, Jonathan, et autres
Publié: (2026)
PilotANN: Memory-Bounded GPU Acceleration for Vector Search
par: Gui, Yuntao, et autres
Publié: (2025)
par: Gui, Yuntao, et autres
Publié: (2025)
MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
par: Hu, Rizhen, et autres
Publié: (2025)
par: Hu, Rizhen, et autres
Publié: (2025)
Collaborative UAVs Multi-task Video Processing Optimization Based on Enhanced Distributed Actor-Critic Networks
par: Rong, Ziqi, et autres
Publié: (2024)
par: Rong, Ziqi, et autres
Publié: (2024)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
par: Guo, Cong, et autres
Publié: (2024)
par: Guo, Cong, et autres
Publié: (2024)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
par: Kong, Jie, et autres
Publié: (2026)
par: Kong, Jie, et autres
Publié: (2026)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
par: Patke, Archit, et autres
Publié: (2025)
par: Patke, Archit, et autres
Publié: (2025)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
par: Chen, Jiu, et autres
Publié: (2026)
par: Chen, Jiu, et autres
Publié: (2026)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
par: Lin, Shouxu, et autres
Publié: (2026)
par: Lin, Shouxu, et autres
Publié: (2026)
Optimizing Federated Learning in the Era of LLMs: Message Quantization and Streaming
par: Xu, Ziyue, et autres
Publié: (2025)
par: Xu, Ziyue, et autres
Publié: (2025)
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
par: Kong, Jie, et autres
Publié: (2025)
par: Kong, Jie, et autres
Publié: (2025)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
par: Hu, Zhisheng, et autres
Publié: (2025)
par: Hu, Zhisheng, et autres
Publié: (2025)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
par: Yang, Shuo, et autres
Publié: (2026)
par: Yang, Shuo, et autres
Publié: (2026)
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
par: Zheng, Haoyu, et autres
Publié: (2025)
par: Zheng, Haoyu, et autres
Publié: (2025)
Heterogeneity-Aware Memory Efficient Federated Learning via Progressive Layer Freezing
par: Yebo, Wu, et autres
Publié: (2024)
par: Yebo, Wu, et autres
Publié: (2024)
Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling
par: Zhang, Xinyi, et autres
Publié: (2024)
par: Zhang, Xinyi, et autres
Publié: (2024)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
par: Xu, Jiale, et autres
Publié: (2025)
par: Xu, Jiale, et autres
Publié: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
par: Zhang, Yuning, et autres
Publié: (2026)
par: Zhang, Yuning, et autres
Publié: (2026)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
par: Pang, Bowen, et autres
Publié: (2025)
par: Pang, Bowen, et autres
Publié: (2025)
Distributed Bilevel Optimization with Dual Pruning for Resource-limited Clients
par: Li, Mingyi, et autres
Publié: (2025)
par: Li, Mingyi, et autres
Publié: (2025)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
par: Duan, Jiangfei, et autres
Publié: (2024)
par: Duan, Jiangfei, et autres
Publié: (2024)
Documents similaires
-
Efficient Long Context Fine-tuning with Chunk Flow
par: Yuan, Xiulong, et autres
Publié: (2025) -
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
par: Sun, Mingyu, et autres
Publié: (2025) -
MemAscend: System Memory Optimization for SSD-Offloaded LLM Fine-Tuning
par: Liaw, Yong-Cheng, et autres
Publié: (2025) -
Analysis and Optimized CXL-Attached Memory Allocation for Long-Context LLM Fine-Tuning
par: Liaw, Yong-Cheng, et autres
Publié: (2025) -
Accelerating Compound LLM Training Workloads with Maestro
par: Yuan, Xiulong, et autres
Publié: (2026)