Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kamath, Aditya K, Krishnamurthy, Arvind, Canini, Marco, Peter, Simon |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
par: Kamath, Aditya K, et autres
Publié: (2024)
par: Kamath, Aditya K, et autres
Publié: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
par: Jo, Myeong Jun
Publié: (2026)
par: Jo, Myeong Jun
Publié: (2026)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
par: Topcu, Burak, et autres
Publié: (2026)
par: Topcu, Burak, et autres
Publié: (2026)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
par: Shakerdargah, Mohammadali, et autres
Publié: (2024)
par: Shakerdargah, Mohammadali, et autres
Publié: (2024)
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
par: Lo, Yun-Chen, et autres
Publié: (2024)
par: Lo, Yun-Chen, et autres
Publié: (2024)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
par: Ganjihal, Sanjeev Rao
Publié: (2026)
par: Ganjihal, Sanjeev Rao
Publié: (2026)
ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
par: Liang, Yuhang, et autres
Publié: (2024)
par: Liang, Yuhang, et autres
Publié: (2024)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
par: Kolluru, Saicharan
Publié: (2025)
par: Kolluru, Saicharan
Publié: (2025)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
par: Penke, Carolin, et autres
Publié: (2025)
par: Penke, Carolin, et autres
Publié: (2025)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
par: Georgiou, Athos
Publié: (2026)
par: Georgiou, Athos
Publié: (2026)
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
par: Xu, Youxuan, et autres
Publié: (2025)
par: Xu, Youxuan, et autres
Publié: (2025)
Addressing tokens dynamic generation, propagation, storage and renewal to secure the GlideinWMS pilot based jobs and system
par: Coimbra, Bruno Moreira, et autres
Publié: (2025)
par: Coimbra, Bruno Moreira, et autres
Publié: (2025)
Kant: An Efficient Unified Scheduling System for Large-Scale AI Clusters
par: Zeng, Lingling, et autres
Publié: (2025)
par: Zeng, Lingling, et autres
Publié: (2025)
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
par: Khan, Kabir, et autres
Publié: (2025)
par: Khan, Kabir, et autres
Publié: (2025)
Flex-MIG: Enabling Distributed Execution on MIG
par: Kim, Myeongsu, et autres
Publié: (2025)
par: Kim, Myeongsu, et autres
Publié: (2025)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
par: Chen, Mu-Chi, et autres
Publié: (2025)
par: Chen, Mu-Chi, et autres
Publié: (2025)
GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
par: Sarker, Yeahia, et autres
Publié: (2026)
par: Sarker, Yeahia, et autres
Publié: (2026)
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
par: Yang, Mingyu, et autres
Publié: (2025)
par: Yang, Mingyu, et autres
Publié: (2025)
Scalable Engine and the Performance of Different LLM Models in a SLURM based HPC architecture
par: Luiz, Anderson de Lima, et autres
Publié: (2025)
par: Luiz, Anderson de Lima, et autres
Publié: (2025)
Combining Serverless and High-Performance Computing Paradigms to support ML Data-Intensive Applications
par: Staylor, Mills, et autres
Publié: (2025)
par: Staylor, Mills, et autres
Publié: (2025)
ZenFlow: Enabling Stall-Free Offloading Training via Asynchronous Updates
par: Lan, Tingfeng, et autres
Publié: (2025)
par: Lan, Tingfeng, et autres
Publié: (2025)
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
par: Chen, Jinfan, et autres
Publié: (2023)
par: Chen, Jinfan, et autres
Publié: (2023)
GREEN-CODE: Learning to Optimize Energy Efficiency in LLM-based Code Generation
par: Ilager, Shashikant, et autres
Publié: (2025)
par: Ilager, Shashikant, et autres
Publié: (2025)
Scalability Evaluation of HPC Multi-GPU Training for ECG-based LLMs
par: Mileski, Dimitar, et autres
Publié: (2025)
par: Mileski, Dimitar, et autres
Publié: (2025)
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
par: Shi, Jinliang, et autres
Publié: (2025)
par: Shi, Jinliang, et autres
Publié: (2025)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
par: Li, Xiangchen, et autres
Publié: (2026)
par: Li, Xiangchen, et autres
Publié: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
par: Li, Xiangchen, et autres
Publié: (2026)
par: Li, Xiangchen, et autres
Publié: (2026)
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
par: Sada, Mohammad Firas, et autres
Publié: (2025)
par: Sada, Mohammad Firas, et autres
Publié: (2025)
StepCache: Step-Level Reuse with Lightweight Verification and Selective Patching for LLM Serving
par: Nouri, Azam
Publié: (2026)
par: Nouri, Azam
Publié: (2026)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
par: Shakeri, Heman, et autres
Publié: (2026)
par: Shakeri, Heman, et autres
Publié: (2026)
Evaluating Large Language Models for Workload Mapping and Scheduling in Heterogeneous HPC Systems
par: Sharma, Aasish Kumar, et autres
Publié: (2025)
par: Sharma, Aasish Kumar, et autres
Publié: (2025)
Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling
par: Luo, Zizhang, et autres
Publié: (2026)
par: Luo, Zizhang, et autres
Publié: (2026)
Accelerating Causal Algorithms for Industrial-scale Data: A Distributed Computing Approach with Ray Framework
par: Verma, Vishal, et autres
Publié: (2024)
par: Verma, Vishal, et autres
Publié: (2024)
Spark-LLM-Eval: A Distributed Framework for Statistically Rigorous Large Language Model Evaluation
par: Mitra, Subhadip
Publié: (2026)
par: Mitra, Subhadip
Publié: (2026)
CRDT-Based Game State Synchronization in Peer-to-Peer VR
par: Dantas, Abel, et autres
Publié: (2025)
par: Dantas, Abel, et autres
Publié: (2025)
Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs
par: Patherya, Kausar, et autres
Publié: (2025)
par: Patherya, Kausar, et autres
Publié: (2025)
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
par: Zehra, Sehar, et autres
Publié: (2025)
par: Zehra, Sehar, et autres
Publié: (2025)
Benchmarking Federated Learning for Throughput Prediction in 5G Live Streaming Applications
par: Dutta, Yuvraj, et autres
Publié: (2025)
par: Dutta, Yuvraj, et autres
Publié: (2025)
DAGER: Exact Gradient Inversion for Large Language Models
par: Petrov, Ivo, et autres
Publié: (2024)
par: Petrov, Ivo, et autres
Publié: (2024)
OPTIMUMP2P: Fast and Reliable Gossiping in P2P Networks
par: Nicolaou, Nicolas, et autres
Publié: (2025)
par: Nicolaou, Nicolas, et autres
Publié: (2025)
Documents similaires
-
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
par: Kamath, Aditya K, et autres
Publié: (2024) -
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
par: Jo, Myeong Jun
Publié: (2026) -
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
par: Topcu, Burak, et autres
Publié: (2026) -
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
par: Shakerdargah, Mohammadali, et autres
Publié: (2024) -
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
par: Lo, Yun-Chen, et autres
Publié: (2024)